For each of the two datasets below

Dataset 1

1979 salaries (in hundreds of Swiss francs) of 7 different professions in 6 cities

Datafile: salaries in LearnEDA package

library(LearnEDAfunctions)
library(tidyverse)
head(salaries)
##   Salary        City Profession
## 1    341   Amsterdam    Teacher
## 2    110      Athens    Teacher
## 3     31     Bangkok    Teacher
## 4    116   Hong_Kong    Teacher
## 5    326 Los_Angeles    Teacher
## 6     89   Singapore    Teacher

♦ Amsterdam \[ LO = 266, F_L = 310, M = 341, F_U = 424,HI = 593 \] \[ step = 171 = 1.5 (114) \,\, {\rm and} \,\, f_s = 114 = 424 - 310 \]
• The inner fences are at \[ 310 - 171 = 139 \,\, {\rm and} \,\, 424 + 171 = 595 \] ♦ Using the same calculation to find the 5-number summaries, fences, and outside values for all the cities, I found:

♠ N = 7 for all cities 

• 1: Amsterdam •

  Depth       Lower        Upper

M: 4.00 341.00
F: 2.50 310.00 424.00

STEP = 171.00
FENCES = 139.00, 595.00
OUTLIERS: None

• 2: Athens •

  Depth       Lower        Upper

M: 4.00 161.00
F: 2.50 117.50 192.00

STEP = 111.75
FENCES = 5.75, 303.75
OUTLIERS: 320

• 3: Bangkok •

  Depth       Lower        Upper

M: 4.00 37.00
F: 2.50 34.50 101.00

STEP = 100.50
FENCES = -66.00, 202.00
OUTLIERS: None

• 4: Hong Kong •

  Depth       Lower        Upper

M: 4.00 116.00
F: 2.50 96.00 159.50

STEP = 95.25
FENCES = 0.75, 254.75
OUTLIERS: None

• 5: Los Angeles •

  Depth       Lower        Upper

M: 4.00 326.00
F: 2.50 308.50 412.00

STEP = 155.25
FENCES = 153.25, 567.25
OUTLIERS: 593.00

• 6: Singapore •

  Depth       Lower        Upper

M: 4.00 89.00
F: 2.50 67.50 97.50

STEP = 45.00
FENCES = 139.00, 595.00
OUTLIERS: 250.00

library(LearnEDAfunctions)
library(tidyverse)
head(salaries)
##   Salary        City Profession
## 1    341   Amsterdam    Teacher
## 2    110      Athens    Teacher
## 3     31     Bangkok    Teacher
## 4    116   Hong_Kong    Teacher
## 5    326 Los_Angeles    Teacher
## 6     89   Singapore    Teacher
####

ggplot(salaries, aes(x = City, y = Salary)) + 
    geom_boxplot() + coord_flip() +
    xlab("City") + ylab("Salary")

spread_level_plot(salaries, Salary, City)

## # A tibble: 6 × 5
##   City            M    df log.M log.df
##   <fct>       <int> <dbl> <dbl>  <dbl>
## 1 Amsterdam     341 114    2.53   2.06
## 2 Athens        161  74.5  2.21   1.87
## 3 Bangkok        37  67    1.57   1.83
## 4 Hong_Kong     116  63.5  2.06   1.80
## 5 Los_Angeles   326 104.   2.51   2.01
## 6 Singapore      89  30    1.95   1.48
####

♦ From the graph => Slope of the line = b = 0.85 ♦ Using the power of the re-expression formula => 1 - b = p • -> 1 - 0.85 = 0.15 and 0.15 is approximately 0 (zero). • -> So, a log transformation should be to stabilize spread.

####

salaries %>% 
    mutate(Reexpressed = log10(Salary)) -> salaries

ggplot(salaries, aes(City, Reexpressed)) + 
    geom_boxplot() + coord_flip()

spread_level_plot(salaries, Reexpressed, City)

## # A tibble: 6 × 5
##   City            M    df log.M log.df
##   <fct>       <dbl> <dbl> <dbl>  <dbl>
## 1 Amsterdam    2.53 0.133 0.404 -0.876
## 2 Athens       2.21 0.214 0.344 -0.669
## 3 Bangkok      1.57 0.457 0.195 -0.340
## 4 Hong_Kong    2.06 0.228 0.315 -0.642
## 5 Los_Angeles  2.51 0.123 0.400 -0.910
## 6 Singapore    1.95 0.171 0.290 -0.768
####

♦ The slope of the line after the log transformation is -2.17, which shows a negative association within this graph along with a dependence between the unequal spreads and levels. I would use the next multiple of one-half to re-express the data without this relationship between level and spread. So, the next next power, p = 0.5, meaning a square-root transformation would be my next attempt

♦ Reexpressed data new 5-number summaries:

♠ N = 7 for all cities 

• 1: Amsterdam •

  Depth       Lower        Upper

M: 4.00 2.53275
F: 2.50 2.49104 2.62395

STEP = 0.19937
FENCES = 2.29167, 2.82331
OUTLIERS: None

• 2: Athens •

  Depth       Lower        Upper

M: 4.00 2.20683
F: 2.50 2.06915 2.28325

STEP = 0.32115
FENCES = 1.74801, 2.60439
OUTLIERS: None

• 3: Bangkok •

  Depth       Lower        Upper

M: 4.00 1.5682
F: 2.50 1.53777 1.9945

STEP = 0.68509
FENCES = 0.85268, 2.06796
OUTLIERS: None

• 4: Hong Kong •

  Depth       Lower        Upper

M: 4.00 2.06446
F: 2.50 1.97359 2.20142

STEP = 0.34174
FENCES = 1.63185, 2.54317
OUTLIERS: None

• 5: Los Angeles •

  Depth       Lower        Upper

M: 4.00 2.51322
F: 2.50 2.48876 2.6118

STEP = 0.18454
FENCES = 2.30423, 2.79635
OUTLIERS: 2.25285

• 6: Singapore •

  Depth       Lower        Upper

M: 4.00 1.94939
F: 2.50 1.81754 1.98831

STEP = 0.25616 FENCES = 1.56138, 2.24447 OUTLIERS: 2.39794

♦ One of the fist differences I noticed after my calculation was the large range of median values from the cities. Bangkok had the lowest median value at 37, followed by Singapore with the next lowest at 89. The highest one was Amsterdam, with a median of 341. Then Los Angeles had the following highest value, at 326. Singapore did have the smallest fourth spread, equaling 30. Although, Amsterdam also had the largest fourth-spread, which was 114. Only three cities had outliers, Athens at 320, Singapore at 250, and LA at 593. Something I noted was that these are all upper outliers, and none of the cities have a single lower outlier. I also observed that that the 593 outlier in Los Angeles was not considered an outlier for Amsterdam due to the difference in the two cities fences. #### Dataset 2

Areas of Important Islands by Continent

Datafile: island.areas in LearnEDA package

head(island.areas)
##    Ocean          Name   Area
## 1 Arctic Axel_Heilberg  16671
## 2 Arctic        Baffin 195928
## 3 Arctic         Banks  27038
## 4 Arctic      Bathurst   6194
## 5 Arctic         Devon  21331
## 6 Arctic     Ellesmere  75767

Compare the island areas of the Artic Ocean, Caribbean Sea, Indian Ocean, Mediterranean Sea and East Indies

• 1: Arctic •

  Depth       Lower        Upper

M: 8.00 16,6671.00
F: 4.50 11,221.00 31,019.00

N = 15 STEP = 29,697.00
FENCES = -18,476.00 amd 60,716.00
OUTLIERS: 75,767 and 83,896 and 195,928

• 2: Caribbean •

  Depth       Lower        Upper

M: 8.00 290.00
F: 4.50 124.00 2,689.00

N= 15 STEP = 3,848.25
FENCES = -3,724.25 and 6,537.75
OUTLIERS: 29,530 and 44,218

• 3: East Indies •

  Depth       Lower        Upper

M: 5.50 51,429.50
F: 3.25 5,672.75 63,975.00

N = 10 STEP = 87,453038
FENCES = -81,780.63 and 151,428.38
OUTLIERS: 165,000 and 280,100

• 4: Indian •

  Depth       Lower        Upper

M: 4.50 844.50
F: 2.75 575.00 8,208.00

N= 8 STEP = 11,449.50
FENCES = -10,874.50 and 19,657.50
OUTLIERS: 25,332 and 22,658

• 5: Mediterranean •

  Depth       Lower        Upper

M: 6.00 1,936.00
F: 3.50 385.50 3,470.50

N = 11 STEP = 4,627.50
FENCES = -4,242.00 and 8098.00
OUTLIERS: 9,262 and 9,822

####

ggplot(island.areas, aes(x = Ocean, y = Area)) + 
    geom_boxplot() + coord_flip() +
    xlab("Ocean") + ylab("Area")

spread_level_plot(island.areas, Area, Ocean)

## # A tibble: 5 × 5
##   Ocean              M     df log.M log.df
##   <fct>          <dbl>  <dbl> <dbl>  <dbl>
## 1 Arctic        16671  19798   4.22   4.30
## 2 Caribbean       290   2566.  2.46   3.41
## 3 East_Indies   21430. 58302.  4.33   4.77
## 4 Indian          844.  7633   2.93   3.88
## 5 Mediterranean  1936   3085   3.29   3.49
####

♦ From the graph => Slope of the line = b = 1.13 ♦ Using the power of the re-expression formula => 1 - b = p • -> 1 - 1.13 = -0.13 and -0.13 is approximately 0 (zero). • -> So, a log transformation should be to stabilize spread.

####

island.areas %>% 
    mutate(Reexpressed = log10(Area)) -> island.areas

ggplot(island.areas, aes(Ocean, Reexpressed)) + 
    geom_boxplot() + coord_flip()

spread_level_plot(island.areas, Reexpressed, Ocean)

## # A tibble: 5 × 5
##   Ocean             M    df log.M   log.df
##   <fct>         <dbl> <dbl> <dbl>    <dbl>
## 1 Arctic         4.22 0.443 0.626 -0.354  
## 2 Caribbean      2.46 1.32  0.391  0.119  
## 3 East_Indies    4.30 1.11  0.634  0.0449 
## 4 Indian         2.92 0.900 0.466 -0.0459 
## 5 Mediterranean  3.29 0.993 0.517 -0.00292
####

♦ The slope of the line after the log transformation is -0.12, which shows a small and minimal dependence between spread and level within this graph. So, I would not attempt to re-express this data anymore because the log transformation converges the slope closer to zero, meaning this transformation alone was sufficient.

♦ Reexpressed data new 5-number summaries:

• 1: Arctic •

  Depth       Lower        Upper

M: 8.00 4.22196
F: 4.50 4.04528 4.48802

N = 15 STEP = 0.66411
FENCES = 3.38117 and 5.15214
OUTLIERS: 5.2921

• 2: Caribbean •

  Depth       Lower        Upper

M: 8.00 2.4624
F: 4.50 2.09252 3.40819

N= 15 STEP = 1.9735
FENCES = 0.11901 and 5.38169
OUTLIERS: None

• 3: East Indies •

  Depth       Lower        Upper

M: 5.50 4.30394
F: 3.25 3.6926 4.80146

N = 10 STEP = 1.6633
FENCES = 2.02931 and 6.46476
OUTLIERS: None

• 4: Indian •

  Depth       Lower        Upper

M: 4.50 2.92183
F: 2.75 2.74958 3.64937

N= 8 STEP = 1.34969
FENCES = 1.39989 and 4.9906
OUTLIERS: 5.35537

• 5: Mediterranean •

  Depth       Lower        Upper

M: 6.00 3.28691
F: 3.50 2.54692 3.54021

N = 11 STEP = 1.48993
FENCES = 1.05698 and 5.03014
OUTLIERS: None

♦ The most notable unsual value I saw were the large outliers within the Oceans before reexpression. East Indies had outliers of 165,00 and 280,100. Arctic had three outliers, 75,767, 83,896 and 195,928. Before the transformation, each ocean also all had at least two outlier values. Then after I applied the transformation, the only Oceans that had any outliers were the Artic and Indian oceans. This shows that that the log transformation I applied helped with the stabilization of the spread. I also noticed the spread and level have a dependence where the areas with larger mean values correlate with larger, wider spreads. As seen from East Indies has the largest median (pre-transformation) of 21,429.50. Along with the largest fourth-spread value of 58,302.25. Similarly, the Caribbean has the smallest median value at 290.00 along the lowest fourth-spread of 2,565.50.