For each of the two datasets below
1979 salaries (in hundreds of Swiss francs) of 7 different professions in 6 cities
Datafile: salaries in LearnEDA package
library(LearnEDAfunctions)
library(tidyverse)
head(salaries)
## Salary City Profession
## 1 341 Amsterdam Teacher
## 2 110 Athens Teacher
## 3 31 Bangkok Teacher
## 4 116 Hong_Kong Teacher
## 5 326 Los_Angeles Teacher
## 6 89 Singapore Teacher
♦ Amsterdam \[
LO = 266, F_L = 310, M = 341, F_U = 424,HI = 593
\] \[
step = 171 = 1.5 (114) \,\, {\rm and} \,\, f_s = 114 = 424 - 310
\]
• The inner fences are at \[
310 - 171 = 139 \,\, {\rm and} \,\, 424 + 171 = 595
\] ♦ Using the same calculation to find the 5-number summaries,
fences, and outside values for all the cities, I found:
♠ N = 7 for all cities
• 1: Amsterdam •
Depth Lower Upper
M: 4.00 341.00
F: 2.50 310.00 424.00
STEP = 171.00
FENCES = 139.00, 595.00
OUTLIERS: None
• 2: Athens •
Depth Lower Upper
M: 4.00 161.00
F: 2.50 117.50 192.00
STEP = 111.75
FENCES = 5.75, 303.75
OUTLIERS: 320
• 3: Bangkok •
Depth Lower Upper
M: 4.00 37.00
F: 2.50 34.50 101.00
STEP = 100.50
FENCES = -66.00, 202.00
OUTLIERS: None
• 4: Hong Kong •
Depth Lower Upper
M: 4.00 116.00
F: 2.50 96.00 159.50
STEP = 95.25
FENCES = 0.75, 254.75
OUTLIERS: None
• 5: Los Angeles •
Depth Lower Upper
M: 4.00 326.00
F: 2.50 308.50 412.00
STEP = 155.25
FENCES = 153.25, 567.25
OUTLIERS: 593.00
• 6: Singapore •
Depth Lower Upper
M: 4.00 89.00
F: 2.50 67.50 97.50
STEP = 45.00
FENCES = 139.00, 595.00
OUTLIERS: 250.00
library(LearnEDAfunctions)
library(tidyverse)
head(salaries)
## Salary City Profession
## 1 341 Amsterdam Teacher
## 2 110 Athens Teacher
## 3 31 Bangkok Teacher
## 4 116 Hong_Kong Teacher
## 5 326 Los_Angeles Teacher
## 6 89 Singapore Teacher
####
ggplot(salaries, aes(x = City, y = Salary)) +
geom_boxplot() + coord_flip() +
xlab("City") + ylab("Salary")
spread_level_plot(salaries, Salary, City)
## # A tibble: 6 × 5
## City M df log.M log.df
## <fct> <int> <dbl> <dbl> <dbl>
## 1 Amsterdam 341 114 2.53 2.06
## 2 Athens 161 74.5 2.21 1.87
## 3 Bangkok 37 67 1.57 1.83
## 4 Hong_Kong 116 63.5 2.06 1.80
## 5 Los_Angeles 326 104. 2.51 2.01
## 6 Singapore 89 30 1.95 1.48
####
♦ From the graph => Slope of the line = b = 0.85 ♦ Using the power of the re-expression formula => 1 - b = p • -> 1 - 0.85 = 0.15 and 0.15 is approximately 0 (zero). • -> So, a log transformation should be to stabilize spread.
####
salaries %>%
mutate(Reexpressed = log10(Salary)) -> salaries
ggplot(salaries, aes(City, Reexpressed)) +
geom_boxplot() + coord_flip()
spread_level_plot(salaries, Reexpressed, City)
## # A tibble: 6 × 5
## City M df log.M log.df
## <fct> <dbl> <dbl> <dbl> <dbl>
## 1 Amsterdam 2.53 0.133 0.404 -0.876
## 2 Athens 2.21 0.214 0.344 -0.669
## 3 Bangkok 1.57 0.457 0.195 -0.340
## 4 Hong_Kong 2.06 0.228 0.315 -0.642
## 5 Los_Angeles 2.51 0.123 0.400 -0.910
## 6 Singapore 1.95 0.171 0.290 -0.768
####
♦ The slope of the line after the log transformation is -2.17, which shows a negative association within this graph along with a dependence between the unequal spreads and levels. I would use the next multiple of one-half to re-express the data without this relationship between level and spread. So, the next next power, p = 0.5, meaning a square-root transformation would be my next attempt
♦ Reexpressed data new 5-number summaries:
♠ N = 7 for all cities
• 1: Amsterdam •
Depth Lower Upper
M: 4.00 2.53275
F: 2.50 2.49104 2.62395
STEP = 0.19937
FENCES = 2.29167, 2.82331
OUTLIERS: None
• 2: Athens •
Depth Lower Upper
M: 4.00 2.20683
F: 2.50 2.06915 2.28325
STEP = 0.32115
FENCES = 1.74801, 2.60439
OUTLIERS: None
• 3: Bangkok •
Depth Lower Upper
M: 4.00 1.5682
F: 2.50 1.53777 1.9945
STEP = 0.68509
FENCES = 0.85268, 2.06796
OUTLIERS: None
• 4: Hong Kong •
Depth Lower Upper
M: 4.00 2.06446
F: 2.50 1.97359 2.20142
STEP = 0.34174
FENCES = 1.63185, 2.54317
OUTLIERS: None
• 5: Los Angeles •
Depth Lower Upper
M: 4.00 2.51322
F: 2.50 2.48876 2.6118
STEP = 0.18454
FENCES = 2.30423, 2.79635
OUTLIERS: 2.25285
• 6: Singapore •
Depth Lower Upper
M: 4.00 1.94939
F: 2.50 1.81754 1.98831
STEP = 0.25616 FENCES = 1.56138, 2.24447 OUTLIERS: 2.39794
♦ One of the fist differences I noticed after my calculation was the large range of median values from the cities. Bangkok had the lowest median value at 37, followed by Singapore with the next lowest at 89. The highest one was Amsterdam, with a median of 341. Then Los Angeles had the following highest value, at 326. Singapore did have the smallest fourth spread, equaling 30. Although, Amsterdam also had the largest fourth-spread, which was 114. Only three cities had outliers, Athens at 320, Singapore at 250, and LA at 593. Something I noted was that these are all upper outliers, and none of the cities have a single lower outlier. I also observed that that the 593 outlier in Los Angeles was not considered an outlier for Amsterdam due to the difference in the two cities fences. #### Dataset 2
Areas of Important Islands by Continent
Datafile: island.areas in LearnEDA package
head(island.areas)
## Ocean Name Area
## 1 Arctic Axel_Heilberg 16671
## 2 Arctic Baffin 195928
## 3 Arctic Banks 27038
## 4 Arctic Bathurst 6194
## 5 Arctic Devon 21331
## 6 Arctic Ellesmere 75767
Compare the island areas of the Artic Ocean, Caribbean Sea, Indian Ocean, Mediterranean Sea and East Indies
• 1: Arctic •
Depth Lower Upper
M: 8.00 16,6671.00
F: 4.50 11,221.00 31,019.00
N = 15 STEP = 29,697.00
FENCES = -18,476.00 amd 60,716.00
OUTLIERS: 75,767 and 83,896 and 195,928
• 2: Caribbean •
Depth Lower Upper
M: 8.00 290.00
F: 4.50 124.00 2,689.00
N= 15 STEP = 3,848.25
FENCES = -3,724.25 and 6,537.75
OUTLIERS: 29,530 and 44,218
• 3: East Indies •
Depth Lower Upper
M: 5.50 51,429.50
F: 3.25 5,672.75 63,975.00
N = 10 STEP = 87,453038
FENCES = -81,780.63 and 151,428.38
OUTLIERS: 165,000 and 280,100
• 4: Indian •
Depth Lower Upper
M: 4.50 844.50
F: 2.75 575.00 8,208.00
N= 8 STEP = 11,449.50
FENCES = -10,874.50 and 19,657.50
OUTLIERS: 25,332 and 22,658
• 5: Mediterranean •
Depth Lower Upper
M: 6.00 1,936.00
F: 3.50 385.50 3,470.50
N = 11 STEP = 4,627.50
FENCES = -4,242.00 and 8098.00
OUTLIERS: 9,262 and 9,822
####
ggplot(island.areas, aes(x = Ocean, y = Area)) +
geom_boxplot() + coord_flip() +
xlab("Ocean") + ylab("Area")
spread_level_plot(island.areas, Area, Ocean)
## # A tibble: 5 × 5
## Ocean M df log.M log.df
## <fct> <dbl> <dbl> <dbl> <dbl>
## 1 Arctic 16671 19798 4.22 4.30
## 2 Caribbean 290 2566. 2.46 3.41
## 3 East_Indies 21430. 58302. 4.33 4.77
## 4 Indian 844. 7633 2.93 3.88
## 5 Mediterranean 1936 3085 3.29 3.49
####
♦ From the graph => Slope of the line = b = 1.13 ♦ Using the power of the re-expression formula => 1 - b = p • -> 1 - 1.13 = -0.13 and -0.13 is approximately 0 (zero). • -> So, a log transformation should be to stabilize spread.
####
island.areas %>%
mutate(Reexpressed = log10(Area)) -> island.areas
ggplot(island.areas, aes(Ocean, Reexpressed)) +
geom_boxplot() + coord_flip()
spread_level_plot(island.areas, Reexpressed, Ocean)
## # A tibble: 5 × 5
## Ocean M df log.M log.df
## <fct> <dbl> <dbl> <dbl> <dbl>
## 1 Arctic 4.22 0.443 0.626 -0.354
## 2 Caribbean 2.46 1.32 0.391 0.119
## 3 East_Indies 4.30 1.11 0.634 0.0449
## 4 Indian 2.92 0.900 0.466 -0.0459
## 5 Mediterranean 3.29 0.993 0.517 -0.00292
####
♦ The slope of the line after the log transformation is -0.12, which shows a small and minimal dependence between spread and level within this graph. So, I would not attempt to re-express this data anymore because the log transformation converges the slope closer to zero, meaning this transformation alone was sufficient.
♦ Reexpressed data new 5-number summaries:
• 1: Arctic •
Depth Lower Upper
M: 8.00 4.22196
F: 4.50 4.04528 4.48802
N = 15 STEP = 0.66411
FENCES = 3.38117 and 5.15214
OUTLIERS: 5.2921
• 2: Caribbean •
Depth Lower Upper
M: 8.00 2.4624
F: 4.50 2.09252 3.40819
N= 15 STEP = 1.9735
FENCES = 0.11901 and 5.38169
OUTLIERS: None
• 3: East Indies •
Depth Lower Upper
M: 5.50 4.30394
F: 3.25 3.6926 4.80146
N = 10 STEP = 1.6633
FENCES = 2.02931 and 6.46476
OUTLIERS: None
• 4: Indian •
Depth Lower Upper
M: 4.50 2.92183
F: 2.75 2.74958 3.64937
N= 8 STEP = 1.34969
FENCES = 1.39989 and 4.9906
OUTLIERS: 5.35537
• 5: Mediterranean •
Depth Lower Upper
M: 6.00 3.28691
F: 3.50 2.54692 3.54021
N = 11 STEP = 1.48993
FENCES = 1.05698 and 5.03014
OUTLIERS: None
♦ The most notable unsual value I saw were the large outliers within the Oceans before reexpression. East Indies had outliers of 165,00 and 280,100. Arctic had three outliers, 75,767, 83,896 and 195,928. Before the transformation, each ocean also all had at least two outlier values. Then after I applied the transformation, the only Oceans that had any outliers were the Artic and Indian oceans. This shows that that the log transformation I applied helped with the stabilization of the spread. I also noticed the spread and level have a dependence where the areas with larger mean values correlate with larger, wider spreads. As seen from East Indies has the largest median (pre-transformation) of 21,429.50. Along with the largest fourth-spread value of 58,302.25. Similarly, the Caribbean has the smallest median value at 290.00 along the lowest fourth-spread of 2,565.50.