Exercise 3.1: Glass Identification

Load and Examine the Data

library(mlbench)
data(Glass)
str(Glass)
## 'data.frame':    214 obs. of  10 variables:
##  $ RI  : num  1.52 1.52 1.52 1.52 1.52 ...
##  $ Na  : num  13.6 13.9 13.5 13.2 13.3 ...
##  $ Mg  : num  4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
##  $ Al  : num  1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
##  $ Si  : num  71.8 72.7 73 72.6 73.1 ...
##  $ K   : num  0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
##  $ Ca  : num  8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
##  $ Ba  : num  0 0 0 0 0 0 0 0 0 0 ...
##  $ Fe  : num  0 0 0 0 0 0.26 0 0 0 0.11 ...
##  $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...

The Glass dataset contains 214 observations, nine numeric predictors, and a categorical outcome called Type. The predictors describe the refractive index and chemical composition of each glass sample.

Exercise 3.1(a): Predictor Distributions

glass_predictors <- Glass[, 1:9]

par(mfrow = c(3, 3), mar = c(4, 4, 2, 1))

for (predictor in names(glass_predictors)) {
  hist(
    glass_predictors[[predictor]],
    main = predictor,
    xlab = predictor,
    col = "blue",
    border = "white"
  )
}

par(mfrow = c(1, 1))

The predictors have noticeably different distributions. K, Ca, Ba, and Fe are strongly right-skewed, with most observations at low values and a few much higher values. RI and Al also show some right skew. Mg has separate clusters near zero and around 3–4, suggesting a distribution with multiple groups. Na is comparatively symmetric, while Si has a longer left tail.

Exercise 3.1(a): Relationships Between Predictors

pairs(
  glass_predictors,
  main = "Relationships Between Glass Predictors",
  pch = 19,
  cex = 0.4,
  col = "green"
)

The scatterplot matrix shows a clear positive relationship between refractive index (RI) and calcium (Ca), and a negative relationship between RI and silicon (Si). Several other predictor pairs show clusters or weaker relationships. K, Ba, and Fe have many observations near zero, with a few points far from the main concentration.

Exercise 3.1(b): Outliers and Skewness

par(mfrow = c(3, 3), mar = c(4, 4, 2, 1))

for (predictor in names(glass_predictors)) {
  boxplot(
    glass_predictors[[predictor]],
    main = predictor,
    ylab = predictor,
    col = "blue"
  )
}

par(mfrow = c(1, 1))

The boxplots flag potential outliers in RI, Na, Al, Si, K, Ca, Ba, and Fe. Mg has no observations beyond the whiskers. RI, Na, Al, Si, and Ca have unusual values at both ends, while K, Ba, and Fe mainly have unusually high values.

Together with the histograms, these plots indicate strong right skew in K, Ca, Ba, and Fe, and some right skew in RI and Al. Si has a longer left tail, while Mg shows separate groups rather than a symmetric distribution.

Ba is concentrated at zero, causing its box and whiskers to collapse near zero and many positive values to be flagged. These flagged observations may represent genuine differences between glass types, so they should be investigated before considering removal.

Exercise 3.1(c): Suggested Transformations

Power transformations, such as the Box–Cox transformation discussed in Chapter 3, could help reduce skewness in suitable predictors. Standard Box–Cox transformations require strictly positive values, so predictors containing zeros need special handling. Transformations may reduce long tails but will not necessarily remove the separate groups seen in Mg or the concentration of zeros in Ba.

Centering and scaling would be useful for classification methods sensitive to predictor scales, such as nearest neighbors and support vector machines. The predictors have substantially different ranges, so standardization would prevent variables with larger scales from dominating distance calculations.

Unusual observations should be investigated before removal because they may represent genuine glass types. Any preprocessing parameters should be estimated from the training data within each resampling split. Classification performance should then be compared to determine whether the transformations improve predictions.

Exercise 3.2: Soybean Data

Load and Examine the Data

data(Soybean, package = "mlbench")
str(Soybean)
## 'data.frame':    683 obs. of  36 variables:
##  $ Class          : Factor w/ 19 levels "2-4-d-injury",..: 11 11 11 11 11 11 11 11 11 11 ...
##  $ date           : Factor w/ 7 levels "0","1","2","3",..: 7 5 4 4 7 6 6 5 7 5 ...
##  $ plant.stand    : Ord.factor w/ 2 levels "0"<"1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ precip         : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
##  $ temp           : Ord.factor w/ 3 levels "0"<"1"<"2": 2 2 2 2 2 2 2 2 2 2 ...
##  $ hail           : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 2 1 1 ...
##  $ crop.hist      : Factor w/ 4 levels "0","1","2","3": 2 3 2 2 3 4 3 2 4 3 ...
##  $ area.dam       : Factor w/ 4 levels "0","1","2","3": 2 1 1 1 1 1 1 1 1 1 ...
##  $ sever          : Factor w/ 3 levels "0","1","2": 2 3 3 3 2 2 2 2 2 3 ...
##  $ seed.tmt       : Factor w/ 3 levels "0","1","2": 1 2 2 1 1 1 2 1 2 1 ...
##  $ germ           : Ord.factor w/ 3 levels "0"<"1"<"2": 1 2 3 2 3 2 1 3 2 3 ...
##  $ plant.growth   : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ leaves         : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ leaf.halo      : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
##  $ leaf.marg      : Factor w/ 3 levels "0","1","2": 3 3 3 3 3 3 3 3 3 3 ...
##  $ leaf.size      : Ord.factor w/ 3 levels "0"<"1"<"2": 3 3 3 3 3 3 3 3 3 3 ...
##  $ leaf.shread    : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ leaf.malf      : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ leaf.mild      : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
##  $ stem           : Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ lodging        : Factor w/ 2 levels "0","1": 2 1 1 1 1 1 2 1 1 1 ...
##  $ stem.cankers   : Factor w/ 4 levels "0","1","2","3": 4 4 4 4 4 4 4 4 4 4 ...
##  $ canker.lesion  : Factor w/ 4 levels "0","1","2","3": 2 2 1 1 2 1 2 2 2 2 ...
##  $ fruiting.bodies: Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 2 ...
##  $ ext.decay      : Factor w/ 3 levels "0","1","2": 2 2 2 2 2 2 2 2 2 2 ...
##  $ mycelium       : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ int.discolor   : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...
##  $ sclerotia      : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ fruit.pods     : Factor w/ 4 levels "0","1","2","3": 1 1 1 1 1 1 1 1 1 1 ...
##  $ fruit.spots    : Factor w/ 4 levels "0","1","2","4": 4 4 4 4 4 4 4 4 4 4 ...
##  $ seed           : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ mold.growth    : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ seed.discolor  : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ seed.size      : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ shriveling     : Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
##  $ roots          : Factor w/ 3 levels "0","1","2": 1 1 1 1 1 1 1 1 1 1 ...

The Soybean dataset contains 683 observations and 35 categorical predictors. The outcome variable, Class, identifies 19 soybean disease classes. The following analysis examines predictor frequencies, predictors with little variation, and missing data.

Exercise 3.2(a): Predictor Frequency Distributions

soybean_predictors <- Soybean[, names(Soybean) != "Class"]

soybean_frequencies <- lapply(
  soybean_predictors,
  table,
  useNA = "ifany"
)

soybean_frequencies
## $date
## 
##    0    1    2    3    4    5    6 <NA> 
##   26   75   93  118  131  149   90    1 
## 
## $plant.stand
## 
##    0    1 <NA> 
##  354  293   36 
## 
## $precip
## 
##    0    1    2 <NA> 
##   74  112  459   38 
## 
## $temp
## 
##    0    1    2 <NA> 
##   80  374  199   30 
## 
## $hail
## 
##    0    1 <NA> 
##  435  127  121 
## 
## $crop.hist
## 
##    0    1    2    3 <NA> 
##   65  165  219  218   16 
## 
## $area.dam
## 
##    0    1    2    3 <NA> 
##  123  227  145  187    1 
## 
## $sever
## 
##    0    1    2 <NA> 
##  195  322   45  121 
## 
## $seed.tmt
## 
##    0    1    2 <NA> 
##  305  222   35  121 
## 
## $germ
## 
##    0    1    2 <NA> 
##  165  213  193  112 
## 
## $plant.growth
## 
##    0    1 <NA> 
##  441  226   16 
## 
## $leaves
## 
##   0   1 
##  77 606 
## 
## $leaf.halo
## 
##    0    1    2 <NA> 
##  221   36  342   84 
## 
## $leaf.marg
## 
##    0    1    2 <NA> 
##  357   21  221   84 
## 
## $leaf.size
## 
##    0    1    2 <NA> 
##   51  327  221   84 
## 
## $leaf.shread
## 
##    0    1 <NA> 
##  487   96  100 
## 
## $leaf.malf
## 
##    0    1 <NA> 
##  554   45   84 
## 
## $leaf.mild
## 
##    0    1    2 <NA> 
##  535   20   20  108 
## 
## $stem
## 
##    0    1 <NA> 
##  296  371   16 
## 
## $lodging
## 
##    0    1 <NA> 
##  520   42  121 
## 
## $stem.cankers
## 
##    0    1    2    3 <NA> 
##  379   39   36  191   38 
## 
## $canker.lesion
## 
##    0    1    2    3 <NA> 
##  320   83  177   65   38 
## 
## $fruiting.bodies
## 
##    0    1 <NA> 
##  473  104  106 
## 
## $ext.decay
## 
##    0    1    2 <NA> 
##  497  135   13   38 
## 
## $mycelium
## 
##    0    1 <NA> 
##  639    6   38 
## 
## $int.discolor
## 
##    0    1    2 <NA> 
##  581   44   20   38 
## 
## $sclerotia
## 
##    0    1 <NA> 
##  625   20   38 
## 
## $fruit.pods
## 
##    0    1    2    3 <NA> 
##  407  130   14   48   84 
## 
## $fruit.spots
## 
##    0    1    2    4 <NA> 
##  345   75   57  100  106 
## 
## $seed
## 
##    0    1 <NA> 
##  476  115   92 
## 
## $mold.growth
## 
##    0    1 <NA> 
##  524   67   92 
## 
## $seed.discolor
## 
##    0    1 <NA> 
##  513   64  106 
## 
## $seed.size
## 
##    0    1 <NA> 
##  532   59   92 
## 
## $shriveling
## 
##    0    1 <NA> 
##  539   38  106 
## 
## $roots
## 
##    0    1    2 <NA> 
##  551   86   15   31

Exercise 3.2(a): Predictors With Little Variation

soybean_variation <- caret::nearZeroVar(
  soybean_predictors,
  saveMetrics = TRUE
)

soybean_variation[
  soybean_variation$zeroVar | soybean_variation$nzv,
  ,
  drop = FALSE
]
##           freqRatio percentUnique zeroVar  nzv
## leaf.mild     26.75     0.4392387   FALSE TRUE
## mycelium     106.50     0.2928258   FALSE TRUE
## sclerotia     31.25     0.2928258   FALSE TRUE

Using the default nearZeroVar() criteria, leaf.mild, mycelium, and sclerotia are flagged as near-zero-variance predictors. Their most frequent categories occur 26.75, 106.50, and 31.25 times as often as their second most frequent categories, respectively. Each predictor also has very few unique categories relative to the number of observations.

These predictors are not constant, but their distributions are highly imbalanced. They are candidates for removal during preprocessing, although their rare categories may help distinguish particular diseases. Their usefulness should therefore be evaluated through resampling before deciding to remove them.

Exercise 3.2(b): Missing Data by Predictor

missing_by_predictor <- data.frame(
  Predictor = names(soybean_predictors),
  MissingCount = colSums(is.na(soybean_predictors)),
  MissingPercent = round(
    100 * colMeans(is.na(soybean_predictors)),
    2
  )
)

missing_by_predictor <- missing_by_predictor[
  order(-missing_by_predictor$MissingCount),
]

knitr::kable(missing_by_predictor, row.names = FALSE)
Predictor MissingCount MissingPercent
hail 121 17.72
sever 121 17.72
seed.tmt 121 17.72
lodging 121 17.72
germ 112 16.40
leaf.mild 108 15.81
fruiting.bodies 106 15.52
fruit.spots 106 15.52
seed.discolor 106 15.52
shriveling 106 15.52
leaf.shread 100 14.64
seed 92 13.47
mold.growth 92 13.47
seed.size 92 13.47
leaf.halo 84 12.30
leaf.marg 84 12.30
leaf.size 84 12.30
leaf.malf 84 12.30
fruit.pods 84 12.30
precip 38 5.56
stem.cankers 38 5.56
canker.lesion 38 5.56
ext.decay 38 5.56
mycelium 38 5.56
int.discolor 38 5.56
sclerotia 38 5.56
plant.stand 36 5.27
roots 31 4.54
temp 30 4.39
crop.hist 16 2.34
plant.growth 16 2.34
stem 16 2.34
date 1 0.15
area.dam 1 0.15
leaves 0 0.00

Exercise 3.2(b): Missing Data by Disease Class

missing_per_row <- rowSums(is.na(soybean_predictors))

missing_by_class <- do.call(
  rbind,
  lapply(split(missing_per_row, Soybean$Class), function(x) {
    data.frame(
      Observations = length(x),
      RowsWithMissing = sum(x > 0),
      PercentWithMissing = round(100 * mean(x > 0), 2),
      AverageMissingPredictors = round(mean(x), 2)
    )
  })
)

missing_by_class$Class <- rownames(missing_by_class)

missing_by_class <- missing_by_class[
  order(-missing_by_class$PercentWithMissing),
  c("Class", "Observations", "RowsWithMissing",
    "PercentWithMissing", "AverageMissingPredictors")
]

knitr::kable(missing_by_class, row.names = FALSE)
Class Observations RowsWithMissing PercentWithMissing AverageMissingPredictors
2-4-d-injury 16 16 100.00 28.12
cyst-nematode 14 14 100.00 24.00
diaporthe-pod-&-stem-blight 15 15 100.00 11.80
herbicide-injury 8 8 100.00 20.00
phytophthora-rot 88 68 77.27 13.80
alternarialeaf-spot 91 0 0.00 0.00
anthracnose 44 0 0.00 0.00
bacterial-blight 20 0 0.00 0.00
bacterial-pustule 20 0 0.00 0.00
brown-spot 92 0 0.00 0.00
brown-stem-rot 44 0 0.00 0.00
charcoal-rot 20 0 0.00 0.00
diaporthe-stem-canker 20 0 0.00 0.00
downy-mildew 20 0 0.00 0.00
frog-eye-leaf-spot 91 0 0.00 0.00
phyllosticta-leaf-spot 20 0 0.00 0.00
powdery-mildew 20 0 0.00 0.00
purple-seed-stain 20 0 0.00 0.00
rhizoctonia-root-rot 20 0 0.00 0.00

Missing values are unevenly distributed across predictors. Hail, sever, seed.tmt, and lodging each have 121 missing values (17.72%), while leaves has none.

Missingness is also concentrated in five disease classes. Every observation in 2-4-d-injury, cyst-nematode, diaporthe-pod-&-stem-blight, and herbicide-injury has at least one missing predictor. For phytophthora-rot, 68 of 88 observations (77.27%) have missing values. The remaining 14 classes have no missing values.

Overall, 121 of 683 observations (17.72%) have at least one missing predictor. This percentage refers to observations, rather than all individual data cells. Removing every incomplete observation would eliminate four entire classes and most observations from a fifth, substantially changing the classification problem.

Exercise 3.2(c): Strategy for Handling Missing Data

I would retain incomplete observations because deleting them would eliminate four disease classes and most observations from a fifth. I would first investigate whether missing entries represent unavailable measurements or symptoms that are not applicable to particular diseases.

For these categorical predictors, an explicit “Missing” category could preserve potentially useful information about missingness. If the entries represent genuinely unavailable measurements, categorical imputation could also be considered. Replacing missing values with the most frequent category is a simple baseline, but it may distort relationships when many values are missing. Numeric median imputation would not be appropriate for unordered categories.

I would compare these approaches using stratified resampling and examine performance for each disease class. Any imputation rules or predictor-removal decisions would be learned within the training portion of each resampling split. The disease class would not be used to impute predictors because it would be unknown when predicting new cases.