Exercises 3.1 and 3.2 from Kuhn & Johnson, Applied Predictive Modeling, Chapter 3 (Data Pre-processing).
data(Glass)
str(Glass)
## 'data.frame': 214 obs. of 10 variables:
## $ RI : num 1.52 1.52 1.52 1.52 1.52 ...
## $ Na : num 13.6 13.9 13.5 13.2 13.3 ...
## $ Mg : num 4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
## $ Al : num 1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
## $ Si : num 71.8 72.7 73 72.6 73.1 ...
## $ K : num 0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
## $ Ca : num 8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
## $ Ba : num 0 0 0 0 0 0 0 0 0 0 ...
## $ Fe : num 0 0 0 0 0 0.26 0 0 0 0.11 ...
## $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...
Glass |>
select(-Type) |>
pivot_longer(everything()) |>
ggplot(aes(value)) +
geom_histogram(bins = 20, fill = "steelblue") +
facet_wrap(~ name, scales = "free") +
labs(title = "Glass predictors: distributions")
glass_cor <- cor(Glass |> select(-Type))
corrplot(glass_cor, method = "ellipse", type = "upper",
addCoef.col = "black", number.cex = 0.7,
title = "Glass predictors: correlation", mar = c(0, 0, 2, 0))
Most predictors (Na, Al, Si, K, Ca, Ba, Fe) are right-skewed with a long tail of larger values; RI and Mg are closer to symmetric, though Mg is bimodal. The strongest relationships are a positive correlation between RI and Ca, and negative correlations between RI/Ca and Si — consistent with how refractive index responds to the oxide composition of the glass.
Glass |>
select(-Type) |>
pivot_longer(everything()) |>
ggplot(aes(x = name, y = value)) +
geom_boxplot(fill = "steelblue", outlier.colour = "red") +
facet_wrap(~ name, scales = "free") +
labs(title = "Glass predictors: boxplots (outliers in red)")
Glass |>
select(-Type) |>
summarise(across(everything(), skewness)) |>
pivot_longer(everything(), names_to = "predictor", values_to = "skewness") |>
arrange(desc(abs(skewness)))
## # A tibble: 9 × 2
## predictor skewness
## <chr> <dbl>
## 1 K 6.46
## 2 Ba 3.37
## 3 Ca 2.02
## 4 Fe 1.73
## 5 RI 1.60
## 6 Mg -1.14
## 7 Al 0.895
## 8 Si -0.720
## 9 Na 0.448
Yes — nearly every predictor has outliers (points well outside the boxplot whiskers), and most are skewed. K, Ba, Ca, Fe, and RI are the most heavily right-skewed (skewness above 1.5), each driven by a small number of high values. Al and Na are only mildly skewed, Mg is left-skewed with a cluster of zeros, and Si is the closest to symmetric.
Given the right-skew and outliers in most predictors, a Box-Cox (or log) transformation would help the heavily skewed, strictly-positive predictors (K, Ba, Ca, Fe, RI, Na) by pulling in their long tails and making outliers less extreme. Box-Cox cannot be applied directly to Mg, since it contains zeros and is left- rather than right-skewed; a different approach (e.g. a spatial-sign transform, or leaving it untransformed) is more appropriate there. Because the predictors are also on different scales and some are correlated, centering and scaling, followed by PCA, would also help a classification model that is sensitive to scale or collinearity.
data(Soybean)
str(Soybean)
## 'data.frame': 683 obs. of 36 variables:
## $ Class : Factor w/ 19 levels "2-4-d-injury",..: 11 11 11 11 11 11 11 11 11 11 ...
## $ date : Factor w/ 7 levels "0","1","2","3",..: 7 5 4 4 7 6 6 5 7 5 ...
## $ plant.stand : Ord.factor w/ 2 levels "0"<"1": 1 1 1 1 1 1 1 1 1 1 ...
## $ precip : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
## $ temp : Ord.factor w/ 3 levels "0"<"1"<"2": 2 2 2 2 2 2 2 2 2 2 ...
## $ hail : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 2 1 1 ...
## $ crop.hist : Factor w/ 4 levels "0","1","2","3": 2 3 2 2 3 4 3 2 4 3 ...
## $ area.dam : Factor w/ 4 levels "0","1","2","3": 2 1 1 1 1 1 1 1 1 1 ...
## $ sever : Factor w/ 3 levels "0","1","2": 2 3 3 3 2 2 2 2 2 3 ...
## $ seed.tmt : Factor w/ 3 levels "0","1","2": 1 2 2 1 1 1 2 1 2 1 ...
## $ germ : Ord.factor w/ 3 levels "0"<"1"<"2": 1 2 3 2 3 2 1 3 2 3 ...
## $ plant.growth : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ leaves : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ leaf.halo : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
## $ leaf.marg : Factor w/ 3 levels "0","1","2": 3 3 3 3 3 3 3 3 3 3 ...
## $ leaf.size : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
## $ leaf.shread : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ leaf.malf : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ leaf.mild : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
## $ stem : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ lodging : Factor w/ 2 levels "0","1": 2 1 1 1 1 1 2 1 1 1 ...
## $ stem.cankers : Factor w/ 4 levels "0","1","2","3": 4 4 4 4 4 4 4 4 4 4 ...
## $ canker.lesion : Factor w/ 4 levels "0","1","2","3": 2 2 1 1 2 1 2 2 2 2 ...
## $ fruiting.bodies: Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ ext.decay : Factor w/ 3 levels "0","1","2": 2 2 2 2 2 2 2 2 2 2 ...
## $ mycelium : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ int.discolor : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
## $ sclerotia : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ fruit.pods : Factor w/ 4 levels "0","1","2","3": 1 1 1 1 1 1 1 1 1 1 ...
## $ fruit.spots : Factor w/ 4 levels "0","1","2","4": 4 4 4 4 4 4 4 4 4 4 ...
## $ seed : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ mold.growth : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ seed.discolor : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ seed.size : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ shriveling : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ roots : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
predictors <- Soybean |> select(-Class)
nzv_tbl <- predictors |>
map_dfr(nzv_summary, .id = "predictor") |>
arrange(desc(freq_ratio))
nzv_tbl
## # A tibble: 35 × 3
## predictor freq_ratio pct_unique
## <chr> <dbl> <dbl>
## 1 mycelium 106. 0.293
## 2 sclerotia 31.2 0.293
## 3 leaf.mild 26.8 0.439
## 4 shriveling 14.2 0.293
## 5 int.discolor 13.2 0.439
## 6 lodging 12.4 0.293
## 7 leaf.malf 12.3 0.293
## 8 seed.size 9.02 0.293
## 9 seed.discolor 8.02 0.293
## 10 leaves 7.87 0.293
## # ℹ 25 more rows
predictors |>
mutate(across(everything(), as.character)) |>
pivot_longer(everything()) |>
filter(!is.na(value)) |>
ggplot(aes(x = value)) +
geom_bar(fill = "steelblue") +
facet_wrap(~ name, scales = "free", ncol = 5) +
theme(axis.text.x = element_text(size = 6)) +
labs(title = "Soybean categorical predictors: level frequencies")
Yes — several predictors have degenerate (near-zero-variance)
distributions: one level is overwhelmingly more common than any
other. By frequency ratio, leaf.mild,
mycelium, and sclerotia stand out,
each with one level covering the vast majority of observations and the
remaining levels barely represented. These contribute almost no
information for discriminating between classes and are candidates for
removal.
Because missingness is concentrated in specific classes rather than scattered at random, two complementary steps make sense:
hail/sever/seed.tmt/lodging
are missing for the same disease classes that give good separation on
other predictors, they can be removed with little loss)."missing" category/level, which preserves the information
that the value was unrecorded rather than discarding those rows.Simply deleting all rows with any missing value is not advisable here — because missingness is concentrated by class, that would delete entire disease classes from the data and bias the model.
sessionInfo()
## R version 4.6.1 (2026-06-24)
## Platform: aarch64-apple-darwin25.4.0
## Running under: macOS Tahoe 26.7
##
## Matrix products: default
## BLAS: /opt/homebrew/Cellar/openblas/0.3.34/lib/libopenblasp-r0.3.34.dylib
## LAPACK: /opt/homebrew/Cellar/r/4.6.1/lib/R/lib/libRlapack.dylib; LAPACK version 3.12.1
##
## locale:
## [1] C.UTF-8/C.UTF-8/C.UTF-8/C/C.UTF-8/C.UTF-8
##
## time zone: America/New_York
## tzcode source: internal
##
## attached base packages:
## [1] stats graphics grDevices utils datasets methods base
##
## other attached packages:
## [1] corrplot_0.95 e1071_1.7-17 tibble_3.3.1 purrr_1.2.2 ggplot2_4.0.3
## [6] tidyr_1.3.2 dplyr_1.2.1 mlbench_2.1-11
##
## loaded via a namespace (and not attached):
## [1] jsonlite_2.0.0 gtable_0.3.6 compiler_4.6.1 tinytex_0.60
## [5] tidyselect_1.2.1 jquerylib_0.1.4 scales_1.4.0 yaml_2.3.12
## [9] fastmap_1.2.0 R6_2.6.1 labeling_0.4.3 generics_0.1.4
## [13] knitr_1.52 bslib_0.12.0 pillar_1.11.1 RColorBrewer_1.1-3
## [17] rlang_1.3.0 utf8_1.2.6 cachem_1.1.0 xfun_0.60
## [21] sass_0.4.10 S7_0.2.2 cli_3.6.6 withr_3.0.3
## [25] magrittr_2.0.5 class_7.3-23 digest_0.6.39 grid_4.6.1
## [29] lifecycle_1.0.5 vctrs_0.7.3 proxy_0.4-29 evaluate_1.0.5
## [33] glue_1.8.1 farver_2.1.2 rmarkdown_2.32 tools_4.6.1
## [37] pkgconfig_2.0.3 htmltools_0.5.9