library(mlbench)
data(Glass)
str(Glass)
## 'data.frame': 214 obs. of 10 variables:
## $ RI : num 1.52 1.52 1.52 1.52 1.52 ...
## $ Na : num 13.6 13.9 13.5 13.2 13.3 ...
## $ Mg : num 4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
## $ Al : num 1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
## $ Si : num 71.8 72.7 73 72.6 73.1 ...
## $ K : num 0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
## $ Ca : num 8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
## $ Ba : num 0 0 0 0 0 0 0 0 0 0 ...
## $ Fe : num 0 0 0 0 0 0.26 0 0 0 0.11 ...
## $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...
The Glass dataset contains 214 observations, nine numeric predictors, and a categorical outcome called Type. The predictors describe the refractive index and chemical composition of each glass sample.
glass_predictors <- Glass[, 1:9]
par(mfrow = c(3, 3), mar = c(4, 4, 2, 1))
for (predictor in names(glass_predictors)) {
hist(
glass_predictors[[predictor]],
main = predictor,
xlab = predictor,
col = "blue",
border = "white"
)
}
par(mfrow = c(1, 1))
The predictors have noticeably different distributions. K, Ca, Ba, and Fe are strongly right-skewed, with most observations at low values and a few much higher values. RI and Al also show some right skew. Mg has separate clusters near zero and around 3–4, suggesting a distribution with multiple groups. Na is comparatively symmetric, while Si has a longer left tail.
pairs(
glass_predictors,
main = "Relationships Between Glass Predictors",
pch = 19,
cex = 0.4,
col = "green"
)
The scatterplot matrix shows a clear positive relationship between
refractive index (RI) and calcium (Ca), and a negative relationship
between RI and silicon (Si). Several other predictor pairs show clusters
or weaker relationships. K, Ba, and Fe have many observations near zero,
with a few points far from the main concentration.
par(mfrow = c(3, 3), mar = c(4, 4, 2, 1))
for (predictor in names(glass_predictors)) {
boxplot(
glass_predictors[[predictor]],
main = predictor,
ylab = predictor,
col = "blue"
)
}
par(mfrow = c(1, 1))
The boxplots flag potential outliers in RI, Na, Al, Si, K, Ca, Ba, and Fe. Mg has no observations beyond the whiskers. RI, Na, Al, Si, and Ca have unusual values at both ends, while K, Ba, and Fe mainly have unusually high values.
Together with the histograms, these plots indicate strong right skew in K, Ca, Ba, and Fe, and some right skew in RI and Al. Si has a longer left tail, while Mg shows separate groups rather than a symmetric distribution.
Ba is concentrated at zero, causing its box and whiskers to collapse near zero and many positive values to be flagged. These flagged observations may represent genuine differences between glass types, so they should be investigated before considering removal.
Power transformations, such as the Box–Cox transformation discussed in Chapter 3, could help reduce skewness in suitable predictors. Standard Box–Cox transformations require strictly positive values, so predictors containing zeros need special handling. Transformations may reduce long tails but will not necessarily remove the separate groups seen in Mg or the concentration of zeros in Ba.
Centering and scaling would be useful for classification methods sensitive to predictor scales, such as nearest neighbors and support vector machines. The predictors have substantially different ranges, so standardization would prevent variables with larger scales from dominating distance calculations.
Unusual observations should be investigated before removal because they may represent genuine glass types. Any preprocessing parameters should be estimated from the training data within each resampling split. Classification performance should then be compared to determine whether the transformations improve predictions.
data(Soybean, package = "mlbench")
str(Soybean)
## 'data.frame': 683 obs. of 36 variables:
## $ Class : Factor w/ 19 levels "2-4-d-injury",..: 11 11 11 11 11 11 11 11 11 11 ...
## $ date : Factor w/ 7 levels "0","1","2","3",..: 7 5 4 4 7 6 6 5 7 5 ...
## $ plant.stand : Ord.factor w/ 2 levels "0"<"1": 1 1 1 1 1 1 1 1 1 1 ...
## $ precip : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
## $ temp : Ord.factor w/ 3 levels "0"<"1"<"2": 2 2 2 2 2 2 2 2 2 2 ...
## $ hail : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 2 1 1 ...
## $ crop.hist : Factor w/ 4 levels "0","1","2","3": 2 3 2 2 3 4 3 2 4 3 ...
## $ area.dam : Factor w/ 4 levels "0","1","2","3": 2 1 1 1 1 1 1 1 1 1 ...
## $ sever : Factor w/ 3 levels "0","1","2": 2 3 3 3 2 2 2 2 2 3 ...
## $ seed.tmt : Factor w/ 3 levels "0","1","2": 1 2 2 1 1 1 2 1 2 1 ...
## $ germ : Ord.factor w/ 3 levels "0"<"1"<"2": 1 2 3 2 3 2 1 3 2 3 ...
## $ plant.growth : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ leaves : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ leaf.halo : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
## $ leaf.marg : Factor w/ 3 levels "0","1","2": 3 3 3 3 3 3 3 3 3 3 ...
## $ leaf.size : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
## $ leaf.shread : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ leaf.malf : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ leaf.mild : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
## $ stem : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ lodging : Factor w/ 2 levels "0","1": 2 1 1 1 1 1 2 1 1 1 ...
## $ stem.cankers : Factor w/ 4 levels "0","1","2","3": 4 4 4 4 4 4 4 4 4 4 ...
## $ canker.lesion : Factor w/ 4 levels "0","1","2","3": 2 2 1 1 2 1 2 2 2 2 ...
## $ fruiting.bodies: Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
## $ ext.decay : Factor w/ 3 levels "0","1","2": 2 2 2 2 2 2 2 2 2 2 ...
## $ mycelium : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ int.discolor : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
## $ sclerotia : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ fruit.pods : Factor w/ 4 levels "0","1","2","3": 1 1 1 1 1 1 1 1 1 1 ...
## $ fruit.spots : Factor w/ 4 levels "0","1","2","4": 4 4 4 4 4 4 4 4 4 4 ...
## $ seed : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ mold.growth : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ seed.discolor : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ seed.size : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ shriveling : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ roots : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
The Soybean dataset contains 683 observations and 35 categorical predictors. The outcome variable, Class, identifies 19 soybean disease classes. The following analysis examines predictor frequencies, predictors with little variation, and missing data.
soybean_predictors <- Soybean[, names(Soybean) != "Class"]
soybean_frequencies <- lapply(
soybean_predictors,
table,
useNA = "ifany"
)
soybean_frequencies
## $date
##
## 0 1 2 3 4 5 6 <NA>
## 26 75 93 118 131 149 90 1
##
## $plant.stand
##
## 0 1 <NA>
## 354 293 36
##
## $precip
##
## 0 1 2 <NA>
## 74 112 459 38
##
## $temp
##
## 0 1 2 <NA>
## 80 374 199 30
##
## $hail
##
## 0 1 <NA>
## 435 127 121
##
## $crop.hist
##
## 0 1 2 3 <NA>
## 65 165 219 218 16
##
## $area.dam
##
## 0 1 2 3 <NA>
## 123 227 145 187 1
##
## $sever
##
## 0 1 2 <NA>
## 195 322 45 121
##
## $seed.tmt
##
## 0 1 2 <NA>
## 305 222 35 121
##
## $germ
##
## 0 1 2 <NA>
## 165 213 193 112
##
## $plant.growth
##
## 0 1 <NA>
## 441 226 16
##
## $leaves
##
## 0 1
## 77 606
##
## $leaf.halo
##
## 0 1 2 <NA>
## 221 36 342 84
##
## $leaf.marg
##
## 0 1 2 <NA>
## 357 21 221 84
##
## $leaf.size
##
## 0 1 2 <NA>
## 51 327 221 84
##
## $leaf.shread
##
## 0 1 <NA>
## 487 96 100
##
## $leaf.malf
##
## 0 1 <NA>
## 554 45 84
##
## $leaf.mild
##
## 0 1 2 <NA>
## 535 20 20 108
##
## $stem
##
## 0 1 <NA>
## 296 371 16
##
## $lodging
##
## 0 1 <NA>
## 520 42 121
##
## $stem.cankers
##
## 0 1 2 3 <NA>
## 379 39 36 191 38
##
## $canker.lesion
##
## 0 1 2 3 <NA>
## 320 83 177 65 38
##
## $fruiting.bodies
##
## 0 1 <NA>
## 473 104 106
##
## $ext.decay
##
## 0 1 2 <NA>
## 497 135 13 38
##
## $mycelium
##
## 0 1 <NA>
## 639 6 38
##
## $int.discolor
##
## 0 1 2 <NA>
## 581 44 20 38
##
## $sclerotia
##
## 0 1 <NA>
## 625 20 38
##
## $fruit.pods
##
## 0 1 2 3 <NA>
## 407 130 14 48 84
##
## $fruit.spots
##
## 0 1 2 4 <NA>
## 345 75 57 100 106
##
## $seed
##
## 0 1 <NA>
## 476 115 92
##
## $mold.growth
##
## 0 1 <NA>
## 524 67 92
##
## $seed.discolor
##
## 0 1 <NA>
## 513 64 106
##
## $seed.size
##
## 0 1 <NA>
## 532 59 92
##
## $shriveling
##
## 0 1 <NA>
## 539 38 106
##
## $roots
##
## 0 1 2 <NA>
## 551 86 15 31
soybean_variation <- caret::nearZeroVar(
soybean_predictors,
saveMetrics = TRUE
)
soybean_variation[
soybean_variation$zeroVar | soybean_variation$nzv,
,
drop = FALSE
]
## freqRatio percentUnique zeroVar nzv
## leaf.mild 26.75 0.4392387 FALSE TRUE
## mycelium 106.50 0.2928258 FALSE TRUE
## sclerotia 31.25 0.2928258 FALSE TRUE
Using the default nearZeroVar() criteria, leaf.mild, mycelium, and sclerotia are flagged as near-zero-variance predictors. Their most frequent categories occur 26.75, 106.50, and 31.25 times as often as their second most frequent categories, respectively. Each predictor also has very few unique categories relative to the number of observations.
These predictors are not constant, but their distributions are highly imbalanced. They are candidates for removal during preprocessing, although their rare categories may help distinguish particular diseases. Their usefulness should therefore be evaluated through resampling before deciding to remove them.
missing_by_predictor <- data.frame(
Predictor = names(soybean_predictors),
MissingCount = colSums(is.na(soybean_predictors)),
MissingPercent = round(
100 * colMeans(is.na(soybean_predictors)),
2
)
)
missing_by_predictor <- missing_by_predictor[
order(-missing_by_predictor$MissingCount),
]
knitr::kable(missing_by_predictor, row.names = FALSE)
| Predictor | MissingCount | MissingPercent |
|---|---|---|
| hail | 121 | 17.72 |
| sever | 121 | 17.72 |
| seed.tmt | 121 | 17.72 |
| lodging | 121 | 17.72 |
| germ | 112 | 16.40 |
| leaf.mild | 108 | 15.81 |
| fruiting.bodies | 106 | 15.52 |
| fruit.spots | 106 | 15.52 |
| seed.discolor | 106 | 15.52 |
| shriveling | 106 | 15.52 |
| leaf.shread | 100 | 14.64 |
| seed | 92 | 13.47 |
| mold.growth | 92 | 13.47 |
| seed.size | 92 | 13.47 |
| leaf.halo | 84 | 12.30 |
| leaf.marg | 84 | 12.30 |
| leaf.size | 84 | 12.30 |
| leaf.malf | 84 | 12.30 |
| fruit.pods | 84 | 12.30 |
| precip | 38 | 5.56 |
| stem.cankers | 38 | 5.56 |
| canker.lesion | 38 | 5.56 |
| ext.decay | 38 | 5.56 |
| mycelium | 38 | 5.56 |
| int.discolor | 38 | 5.56 |
| sclerotia | 38 | 5.56 |
| plant.stand | 36 | 5.27 |
| roots | 31 | 4.54 |
| temp | 30 | 4.39 |
| crop.hist | 16 | 2.34 |
| plant.growth | 16 | 2.34 |
| stem | 16 | 2.34 |
| date | 1 | 0.15 |
| area.dam | 1 | 0.15 |
| leaves | 0 | 0.00 |
missing_per_row <- rowSums(is.na(soybean_predictors))
missing_by_class <- do.call(
rbind,
lapply(split(missing_per_row, Soybean$Class), function(x) {
data.frame(
Observations = length(x),
RowsWithMissing = sum(x > 0),
PercentWithMissing = round(100 * mean(x > 0), 2),
AverageMissingPredictors = round(mean(x), 2)
)
})
)
missing_by_class$Class <- rownames(missing_by_class)
missing_by_class <- missing_by_class[
order(-missing_by_class$PercentWithMissing),
c("Class", "Observations", "RowsWithMissing",
"PercentWithMissing", "AverageMissingPredictors")
]
knitr::kable(missing_by_class, row.names = FALSE)
| Class | Observations | RowsWithMissing | PercentWithMissing | AverageMissingPredictors |
|---|---|---|---|---|
| 2-4-d-injury | 16 | 16 | 100.00 | 28.12 |
| cyst-nematode | 14 | 14 | 100.00 | 24.00 |
| diaporthe-pod-&-stem-blight | 15 | 15 | 100.00 | 11.80 |
| herbicide-injury | 8 | 8 | 100.00 | 20.00 |
| phytophthora-rot | 88 | 68 | 77.27 | 13.80 |
| alternarialeaf-spot | 91 | 0 | 0.00 | 0.00 |
| anthracnose | 44 | 0 | 0.00 | 0.00 |
| bacterial-blight | 20 | 0 | 0.00 | 0.00 |
| bacterial-pustule | 20 | 0 | 0.00 | 0.00 |
| brown-spot | 92 | 0 | 0.00 | 0.00 |
| brown-stem-rot | 44 | 0 | 0.00 | 0.00 |
| charcoal-rot | 20 | 0 | 0.00 | 0.00 |
| diaporthe-stem-canker | 20 | 0 | 0.00 | 0.00 |
| downy-mildew | 20 | 0 | 0.00 | 0.00 |
| frog-eye-leaf-spot | 91 | 0 | 0.00 | 0.00 |
| phyllosticta-leaf-spot | 20 | 0 | 0.00 | 0.00 |
| powdery-mildew | 20 | 0 | 0.00 | 0.00 |
| purple-seed-stain | 20 | 0 | 0.00 | 0.00 |
| rhizoctonia-root-rot | 20 | 0 | 0.00 | 0.00 |
Missing values are unevenly distributed across predictors. Hail, sever, seed.tmt, and lodging each have 121 missing values (17.72%), while leaves has none.
Missingness is also concentrated in five disease classes. Every observation in 2-4-d-injury, cyst-nematode, diaporthe-pod-&-stem-blight, and herbicide-injury has at least one missing predictor. For phytophthora-rot, 68 of 88 observations (77.27%) have missing values. The remaining 14 classes have no missing values.
Overall, 121 of 683 observations (17.72%) have at least one missing predictor. This percentage refers to observations, rather than all individual data cells. Removing every incomplete observation would eliminate four entire classes and most observations from a fifth, substantially changing the classification problem.
I would retain incomplete observations because deleting them would eliminate four disease classes and most observations from a fifth. I would first investigate whether missing entries represent unavailable measurements or symptoms that are not applicable to particular diseases.
For these categorical predictors, an explicit “Missing” category could preserve potentially useful information about missingness. If the entries represent genuinely unavailable measurements, categorical imputation could also be considered. Replacing missing values with the most frequent category is a simple baseline, but it may distort relationships when many values are missing. Numeric median imputation would not be appropriate for unordered categories.
I would compare these approaches using stratified resampling and examine performance for each disease class. Any imputation rules or predictor-removal decisions would be learned within the training portion of each resampling split. The disease class would not be used to impute predictors because it would be unknown when predicting new cases.