Exercises 3.1 and 3.2 from Kuhn & Johnson, Applied Predictive Modeling, Chapter 3 (Data Pre-processing).


Exercise 3.1 — Glass identification data

data(Glass)
str(Glass)
## 'data.frame':    214 obs. of  10 variables:
##  $ RI  : num  1.52 1.52 1.52 1.52 1.52 ...
##  $ Na  : num  13.6 13.9 13.5 13.2 13.3 ...
##  $ Mg  : num  4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
##  $ Al  : num  1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
##  $ Si  : num  71.8 72.7 73 72.6 73.1 ...
##  $ K   : num  0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
##  $ Ca  : num  8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
##  $ Ba  : num  0 0 0 0 0 0 0 0 0 0 ...
##  $ Fe  : num  0 0 0 0 0 0.26 0 0 0 0.11 ...
##  $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...

(a) Explore the predictors: distributions and relationships

Glass |>
  select(-Type) |>
  pivot_longer(everything()) |>
  ggplot(aes(value)) +
  geom_histogram(bins = 20, fill = "steelblue") +
  facet_wrap(~ name, scales = "free") +
  labs(title = "Glass predictors: distributions")

glass_cor <- cor(Glass |> select(-Type))
corrplot(glass_cor, method = "ellipse", type = "upper",
         addCoef.col = "black", number.cex = 0.7,
         title = "Glass predictors: correlation", mar = c(0, 0, 2, 0))

Most predictors (Na, Al, Si, K, Ca, Ba, Fe) are right-skewed with a long tail of larger values; RI and Mg are closer to symmetric, though Mg is bimodal. The strongest relationships are a positive correlation between RI and Ca, and negative correlations between RI/Ca and Si — consistent with how refractive index responds to the oxide composition of the glass.

(b) Outliers and skewness

Glass |>
  select(-Type) |>
  pivot_longer(everything()) |>
  ggplot(aes(x = name, y = value)) +
  geom_boxplot(fill = "steelblue", outlier.colour = "red") +
  facet_wrap(~ name, scales = "free") +
  labs(title = "Glass predictors: boxplots (outliers in red)")

Glass |>
  select(-Type) |>
  summarise(across(everything(), skewness)) |>
  pivot_longer(everything(), names_to = "predictor", values_to = "skewness") |>
  arrange(desc(abs(skewness)))
## # A tibble: 9 × 2
##   predictor skewness
##   <chr>        <dbl>
## 1 K            6.46 
## 2 Ba           3.37 
## 3 Ca           2.02 
## 4 Fe           1.73 
## 5 RI           1.60 
## 6 Mg          -1.14 
## 7 Al           0.895
## 8 Si          -0.720
## 9 Na           0.448

Yes — nearly every predictor has outliers (points well outside the boxplot whiskers), and most are skewed. K, Ba, Ca, Fe, and RI are the most heavily right-skewed (skewness above 1.5), each driven by a small number of high values. Al and Na are only mildly skewed, Mg is left-skewed with a cluster of zeros, and Si is the closest to symmetric.

(c) Transformations that might help

Given the right-skew and outliers in most predictors, a Box-Cox (or log) transformation would help the heavily skewed, strictly-positive predictors (K, Ba, Ca, Fe, RI, Na) by pulling in their long tails and making outliers less extreme. Box-Cox cannot be applied directly to Mg, since it contains zeros and is left- rather than right-skewed; a different approach (e.g. a spatial-sign transform, or leaving it untransformed) is more appropriate there. Because the predictors are also on different scales and some are correlated, centering and scaling, followed by PCA, would also help a classification model that is sensitive to scale or collinearity.


Exercise 3.2 — Soybean data

data(Soybean)
str(Soybean)
## 'data.frame':    683 obs. of  36 variables:
##  $ Class          : Factor w/ 19 levels "2-4-d-injury",..: 11 11 11 11 11 11 11 11 11 11 ...
##  $ date           : Factor w/ 7 levels "0","1","2","3",..: 7 5 4 4 7 6 6 5 7 5 ...
##  $ plant.stand    : Ord.factor w/ 2 levels "0"<"1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ precip         : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
##  $ temp           : Ord.factor w/ 3 levels "0"<"1"<"2": 2 2 2 2 2 2 2 2 2 2 ...
##  $ hail           : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 2 1 1 ...
##  $ crop.hist      : Factor w/ 4 levels "0","1","2","3": 2 3 2 2 3 4 3 2 4 3 ...
##  $ area.dam       : Factor w/ 4 levels "0","1","2","3": 2 1 1 1 1 1 1 1 1 1 ...
##  $ sever          : Factor w/ 3 levels "0","1","2": 2 3 3 3 2 2 2 2 2 3 ...
##  $ seed.tmt       : Factor w/ 3 levels "0","1","2": 1 2 2 1 1 1 2 1 2 1 ...
##  $ germ           : Ord.factor w/ 3 levels "0"<"1"<"2": 1 2 3 2 3 2 1 3 2 3 ...
##  $ plant.growth   : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ leaves         : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ leaf.halo      : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
##  $ leaf.marg      : Factor w/ 3 levels "0","1","2": 3 3 3 3 3 3 3 3 3 3 ...
##  $ leaf.size      : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
##  $ leaf.shread    : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ leaf.malf      : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ leaf.mild      : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
##  $ stem           : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ lodging        : Factor w/ 2 levels "0","1": 2 1 1 1 1 1 2 1 1 1 ...
##  $ stem.cankers   : Factor w/ 4 levels "0","1","2","3": 4 4 4 4 4 4 4 4 4 4 ...
##  $ canker.lesion  : Factor w/ 4 levels "0","1","2","3": 2 2 1 1 2 1 2 2 2 2 ...
##  $ fruiting.bodies: Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ ext.decay      : Factor w/ 3 levels "0","1","2": 2 2 2 2 2 2 2 2 2 2 ...
##  $ mycelium       : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ int.discolor   : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
##  $ sclerotia      : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ fruit.pods     : Factor w/ 4 levels "0","1","2","3": 1 1 1 1 1 1 1 1 1 1 ...
##  $ fruit.spots    : Factor w/ 4 levels "0","1","2","4": 4 4 4 4 4 4 4 4 4 4 ...
##  $ seed           : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ mold.growth    : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ seed.discolor  : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ seed.size      : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ shriveling     : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ roots          : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...

(a) Frequency distributions of the categorical predictors

predictors <- Soybean |> select(-Class)

nzv_tbl <- predictors |>
  map_dfr(nzv_summary, .id = "predictor") |>
  arrange(desc(freq_ratio))

nzv_tbl
## # A tibble: 35 × 3
##    predictor     freq_ratio pct_unique
##    <chr>              <dbl>      <dbl>
##  1 mycelium          106.        0.293
##  2 sclerotia          31.2       0.293
##  3 leaf.mild          26.8       0.439
##  4 shriveling         14.2       0.293
##  5 int.discolor       13.2       0.439
##  6 lodging            12.4       0.293
##  7 leaf.malf          12.3       0.293
##  8 seed.size           9.02      0.293
##  9 seed.discolor       8.02      0.293
## 10 leaves              7.87      0.293
## # ℹ 25 more rows
predictors |>
  mutate(across(everything(), as.character)) |>
  pivot_longer(everything()) |>
  filter(!is.na(value)) |>
  ggplot(aes(x = value)) +
  geom_bar(fill = "steelblue") +
  facet_wrap(~ name, scales = "free", ncol = 5) +
  theme(axis.text.x = element_text(size = 6)) +
  labs(title = "Soybean categorical predictors: level frequencies")

Yes — several predictors have degenerate (near-zero-variance) distributions: one level is overwhelmingly more common than any other. By frequency ratio, leaf.mild, mycelium, and sclerotia stand out, each with one level covering the vast majority of observations and the remaining levels barely represented. These contribute almost no information for discriminating between classes and are candidates for removal.

(c) A strategy for handling the missing data

Because missingness is concentrated in specific classes rather than scattered at random, two complementary steps make sense:

  1. Drop predictors that are missing for almost an entire class and provide little information elsewhere (e.g. if hail/sever/seed.tmt/lodging are missing for the same disease classes that give good separation on other predictors, they can be removed with little loss).
  2. Impute the rest, since missingness is informative but the predictors themselves still carry signal: for a categorical predictor like these, a simple approach is to impute with the most frequent level within each class (since missingness clusters by class), or to add an explicit "missing" category/level, which preserves the information that the value was unrecorded rather than discarding those rows.

Simply deleting all rows with any missing value is not advisable here — because missingness is concentrated by class, that would delete entire disease classes from the data and bias the model.


Session info

sessionInfo()
## R version 4.6.1 (2026-06-24)
## Platform: aarch64-apple-darwin25.4.0
## Running under: macOS Tahoe 26.7
## 
## Matrix products: default
## BLAS:   /opt/homebrew/Cellar/openblas/0.3.34/lib/libopenblasp-r0.3.34.dylib 
## LAPACK: /opt/homebrew/Cellar/r/4.6.1/lib/R/lib/libRlapack.dylib;  LAPACK version 3.12.1
## 
## locale:
## [1] C.UTF-8/C.UTF-8/C.UTF-8/C/C.UTF-8/C.UTF-8
## 
## time zone: America/New_York
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] corrplot_0.95  e1071_1.7-17   tibble_3.3.1   purrr_1.2.2    ggplot2_4.0.3 
## [6] tidyr_1.3.2    dplyr_1.2.1    mlbench_2.1-11
## 
## loaded via a namespace (and not attached):
##  [1] jsonlite_2.0.0     gtable_0.3.6       compiler_4.6.1     tinytex_0.60      
##  [5] tidyselect_1.2.1   jquerylib_0.1.4    scales_1.4.0       yaml_2.3.12       
##  [9] fastmap_1.2.0      R6_2.6.1           labeling_0.4.3     generics_0.1.4    
## [13] knitr_1.52         bslib_0.12.0       pillar_1.11.1      RColorBrewer_1.1-3
## [17] rlang_1.3.0        utf8_1.2.6         cachem_1.1.0       xfun_0.60         
## [21] sass_0.4.10        S7_0.2.2           cli_3.6.6          withr_3.0.3       
## [25] magrittr_2.0.5     class_7.3-23       digest_0.6.39      grid_4.6.1        
## [29] lifecycle_1.0.5    vctrs_0.7.3        proxy_0.4-29       evaluate_1.0.5    
## [33] glue_1.8.1         farver_2.1.2       rmarkdown_2.32     tools_4.6.1       
## [37] pkgconfig_2.0.3    htmltools_0.5.9