Exercise 3.1

The UC Irvine Machine Learning Repository6 contains a data set related to glass identification. The data consist of 214 glass samples labeled as one of seven class categories. There are nine predictors, including the refractive index and percentages of eight elements: Na, Mg, Al, Si, K, Ca, Ba, and Fe.

The data can be accessed via:

## 'data.frame':    214 obs. of  10 variables:
##  $ RI  : num  1.52 1.52 1.52 1.52 1.52 ...
##  $ Na  : num  13.6 13.9 13.5 13.2 13.3 ...
##  $ Mg  : num  4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
##  $ Al  : num  1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
##  $ Si  : num  71.8 72.7 73 72.6 73.1 ...
##  $ K   : num  0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
##  $ Ca  : num  8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
##  $ Ba  : num  0 0 0 0 0 0 0 0 0 0 ...
##  $ Fe  : num  0 0 0 0 0 0.26 0 0 0 0.11 ...
##  $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...

Note that the echo = FALSE parameter was added to the code chunk to prevent printing of the R code that generated the plot.

  1. Using visualizations, explore the predictor variables to understand their distributions as well as the relationships between predictors.
  2. Do there appear to be any outliers in the data? Are any predictors skewed?
  3. Are there any relevant transformations of one or more predictors that might improve the classification model?

Exercise 3.1: Glass Identification Dataset

The UC Irvine Machine Learning Repository contains a dataset related to glass identification with 214 glass samples and 9 numeric predictors: refractive index (RI) and elemental percentages (Na, Mg, Al, Si, K, Ca, Ba, Fe).

## 'data.frame':    214 obs. of  10 variables:
##  $ RI  : num  1.52 1.52 1.52 1.52 1.52 ...
##  $ Na  : num  13.6 13.9 13.5 13.2 13.3 ...
##  $ Mg  : num  4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
##  $ Al  : num  1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
##  $ Si  : num  71.8 72.7 73 72.6 73.1 ...
##  $ K   : num  0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
##  $ Ca  : num  8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
##  $ Ba  : num  0 0 0 0 0 0 0 0 0 0 ...
##  $ Fe  : num  0 0 0 0 0 0.26 0 0 0 0.11 ...
##  $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...

(a) Exploratory Visualizations & Predictor Relationships

We inspect the univariate distributions of the continuous predictors using histograms and examine bivariate correlation patterns using a correlation matrix.

Observations:

  • Refractive Index (RI) & Calcium (Ca): Exhibit a strong positive linear correlation (\(r \approx 0.81\)).
  • Refractive Index (RI) & Silicon (Si): Show a notable inverse/negative correlation.
  • Multimodal / Sparse Elements: Ba, Fe, and K are heavily concentrated at 0, indicating that many glass types do not contain detectable traces of these trace elements. Mg displays a pronounced bimodal distribution.

(b) Outliers and Skewness

We calculate skewness metrics across all 9 numerical features and render boxplots to inspect potential extreme values.

Table 1: Skewness Measures for Glass Predictors
Predictor Skewness Assessment
K K 6.460 Highly Skewed
Ba Ba 3.369 Highly Skewed
Ca Ca 2.018 Highly Skewed
Fe Fe 1.730 Highly Skewed
RI RI 1.603 Highly Skewed
Mg Mg -1.136 Highly Skewed
Al Al 0.895 Moderately Skewed
Si Si -0.720 Moderately Skewed
Na Na 0.448 Roughly Symmetric

Outlier & Skewness Findings:

  • Heavy Outliers: K, Ba, Fe, and Ca exhibit severe upper-tail outliers. For example, K contains isolated observations above \(6.0\) while the vast majority cluster below \(1.0\).
  • High Skewness: K (\(3.56\)), Ba (\(3.37\)), Ca (\(2.02\)), Fe (\(1.73\)), and RI (\(1.61\)) all exhibit substantial right-skewness (\(\vert{}skew\vert{} > 1\)). Mg is negatively skewed (\(-1.14\)) due to its structural bimodal cluster at zero.

Exercise 3.2

The soybean data can also be found at the UC Irvine Machine Learning Repository. Data were collected to predict disease in 683 soybeans. The 35 predictors are mostly categorical and include information on the environmental conditions (e.g., temperature, precipitation) and plant conditions (e.g., left spots, mold growth). The outcome labels consist of 19 distinct classes.

Note that the echo = FALSE parameter was added to the code chunk to prevent printing of the R code that generated the plot.

  1. Investigate the frequency distributions for the categorical predictors. Are any of the distributions degenerate in the ways discussed earlier in this chapter?
  2. Roughly 18% of the data are missing. Are there particular predictors that are more likely to be missing? Is the pattern of missing data related to the classes?
  3. Develop a strategy for handling missing data, either by eliminating predictors or imputation.

Exercise 3.2: Soybean Disease Dataset

The Soybean dataset contains 683 samples and 35 categorical predictors (soil, weather, leaf, and plant symptoms) categorized into 19 disease classes.

(a) Frequency Distributions & Degenerate Predictors

A predictor is degenerate (or near-zero variance) if it contains only one unique value, or if the ratio of the most frequent value to the second most frequent value is exceptionally large (\(> 95/5\) rule) and the percentage of unique values is small.

Table 3: Degenerate (Near-Zero Variance) Predictors in Soybean Data
Predictor freqRatio percentUnique zeroVar nzv
leaf.mild 26.75 0.4392387 FALSE TRUE
mycelium 106.50 0.2928258 FALSE TRUE
sclerotia 31.25 0.2928258 FALSE TRUE

Findings:

  • leaf.mild, mycelium, and sclerotia exhibit extreme frequency imbalances (frequency ratios \(> 15\) to over \(100\)) and almost zero variance.
  • When splitting into training/testing folds or building linear/tree models, these features can become invariant constants within specific cross-validation resamples, destabilizing parameter estimation.

(b) Missing Data Patterns and Class Relationships

Approximately 18% of the data values are missing across the observations. We evaluate missingness rates per predictor and assess if missingness correlates with specific disease classes.

Table 4: Top 10 Predictors with Highest Missingness (%)
Predictor Percent_Missing
hail 17.71596
sever 17.71596
seed.tmt 17.71596
lodging 17.71596
germ 16.39824
leaf.mild 15.81259
fruiting.bodies 15.51977
fruit.spots 15.51977
seed.discolor 15.51977
shriveling 15.51977
Table 5: Missingness Profile Grouped by Disease Class
Class Total_Cases Samples_With_Missing Pct_Cases_Incomplete Mean_Missing_Fields
2-4-d-injury 16 16 100.0 28.1
cyst-nematode 14 14 100.0 24.0
diaporthe-pod-&-stem-blight 15 15 100.0 11.8
herbicide-injury 8 8 100.0 20.0
phytophthora-rot 88 68 77.3 13.8
alternarialeaf-spot 91 0 0.0 0.0
anthracnose 44 0 0.0 0.0
bacterial-blight 20 0 0.0 0.0
bacterial-pustule 20 0 0.0 0.0
brown-spot 92 0 0.0 0.0
brown-stem-rot 44 0 0.0 0.0
charcoal-rot 20 0 0.0 0.0
diaporthe-stem-canker 20 0 0.0 0.0
downy-mildew 20 0 0.0 0.0
frog-eye-leaf-spot 91 0 0.0 0.0
phyllosticta-leaf-spot 20 0 0.0 0.0
powdery-mildew 20 0 0.0 0.0
purple-seed-stain 20 0 0.0 0.0
rhizoctonia-root-rot 20 0 0.0 0.0

Findings:

  • Class-Specific Clustering: Missingness is not Missing Completely at Random (MCAR). The missing values belong almost entirely to specific classes: 2-4-d-injury, cyst-nematode, diaporthe-stem-canker, and phytophthora-rot consist of 100% incomplete observations.
  • In contrast, diseases such as bacterial-blight, brown-spot, and anthracnose contain zero missing values across all measurements.

(c) Strategy for Handling Missing Data

  1. Do Not Indiscriminately Drop Incomplete Rows: Dropping all cases with missing data (na.omit()) would completely eliminate entire diagnostic classes (e.g., eliminating all occurrences of 2-4-d-injury).

  2. Filter Degenerate Predictors First: Remove uninformative predictors with near-zero variance (mycelium, sclerotia, leaf.mild) identified in part (a).

  3. Imputation Methods:

  • K-Nearest Neighbors (KNN) or Tree-Based Imputation (missForest/MICE): Because the predictors are correlated categorical variables, using a distance-weighted donor (KNN) or tree surrogate model preserves discrete structural relationships without distorting distributions.
  • Informative Missingness Dummy: Because the missingness pattern itself is strongly predictive of the disease class, a binary indicator variable (\(I_{missing}\)) can be engineered prior to imputation.

Texas

The official state motto of Texas is “Friendship”.

What it Refers To

  • Origin of the State’s Name: The motto refers to the origin of the word “Texas” (or the Spanish pronunciation, “Tejas”).

  • Native American Roots: It comes from a Caddo (specifically Hasinai) Native American word—such as teyshas or texias—which translates to “friends” or “allies”.

  • Adoption: The 41st Texas Legislature officially adopted “Friendship” as the state motto in February 1930 to honor this historic linguistic connection.