The UC Irvine Machine Learning Repository6 contains a data set related to glass identification. The data consist of 214 glass samples labeled as one of seven class categories. There are nine predictors, including the refractive index and percentages of eight elements: Na, Mg, Al, Si, K, Ca, Ba, and Fe.
The data can be accessed via:
## 'data.frame': 214 obs. of 10 variables:
## $ RI : num 1.52 1.52 1.52 1.52 1.52 ...
## $ Na : num 13.6 13.9 13.5 13.2 13.3 ...
## $ Mg : num 4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
## $ Al : num 1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
## $ Si : num 71.8 72.7 73 72.6 73.1 ...
## $ K : num 0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
## $ Ca : num 8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
## $ Ba : num 0 0 0 0 0 0 0 0 0 0 ...
## $ Fe : num 0 0 0 0 0 0.26 0 0 0 0.11 ...
## $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...
Note that the echo = FALSE parameter was added to the
code chunk to prevent printing of the R code that generated the
plot.
The UC Irvine Machine Learning Repository contains a dataset related
to glass identification with 214 glass samples and 9 numeric predictors:
refractive index (RI) and elemental percentages
(Na, Mg, Al, Si,
K, Ca, Ba, Fe).
## 'data.frame': 214 obs. of 10 variables:
## $ RI : num 1.52 1.52 1.52 1.52 1.52 ...
## $ Na : num 13.6 13.9 13.5 13.2 13.3 ...
## $ Mg : num 4.49 3.6 3.55 3.69 3.62 3.61 3.6 3.61 3.58 3.6 ...
## $ Al : num 1.1 1.36 1.54 1.29 1.24 1.62 1.14 1.05 1.37 1.36 ...
## $ Si : num 71.8 72.7 73 72.6 73.1 ...
## $ K : num 0.06 0.48 0.39 0.57 0.55 0.64 0.58 0.57 0.56 0.57 ...
## $ Ca : num 8.75 7.83 7.78 8.22 8.07 8.07 8.17 8.24 8.3 8.4 ...
## $ Ba : num 0 0 0 0 0 0 0 0 0 0 ...
## $ Fe : num 0 0 0 0 0 0.26 0 0 0 0.11 ...
## $ Type: Factor w/ 6 levels "1","2","3","5",..: 1 1 1 1 1 1 1 1 1 1 ...
We inspect the univariate distributions of the continuous predictors using histograms and examine bivariate correlation patterns using a correlation matrix.
RI) & Calcium
(Ca): Exhibit a strong positive linear correlation
(\(r \approx 0.81\)).RI) & Silicon
(Si): Show a notable inverse/negative
correlation.Ba,
Fe, and K are heavily concentrated at 0,
indicating that many glass types do not contain detectable traces of
these trace elements. Mg displays a pronounced bimodal
distribution.We calculate skewness metrics across all 9 numerical features and render boxplots to inspect potential extreme values.
| Predictor | Skewness | Assessment | |
|---|---|---|---|
| K | K | 6.460 | Highly Skewed |
| Ba | Ba | 3.369 | Highly Skewed |
| Ca | Ca | 2.018 | Highly Skewed |
| Fe | Fe | 1.730 | Highly Skewed |
| RI | RI | 1.603 | Highly Skewed |
| Mg | Mg | -1.136 | Highly Skewed |
| Al | Al | 0.895 | Moderately Skewed |
| Si | Si | -0.720 | Moderately Skewed |
| Na | Na | 0.448 | Roughly Symmetric |
K, Ba,
Fe, and Ca exhibit severe upper-tail outliers.
For example, K contains isolated observations above \(6.0\) while the vast majority cluster below
\(1.0\).K (\(3.56\)), Ba (\(3.37\)), Ca (\(2.02\)), Fe (\(1.73\)), and RI (\(1.61\)) all exhibit substantial
right-skewness (\(\vert{}skew\vert{} >
1\)). Mg is negatively skewed (\(-1.14\)) due to its structural bimodal
cluster at zero.Due to strong right-skewness, large scale disparities, and high collinearity:
Si and Ca span much larger ranges than trace
minerals (Fe, Ba). Standardizing them to \(\mu = 0, \sigma = 1\) is necessary for
distance-based or regularized models.RI, Na,
Al, Si, Ca) benefit from standard
Box-Cox. Predictors with zero values (Ba, Fe,
K, Mg) require either the
Yeo-Johnson transformation or spatial sign
filtering.RI and Ca are strongly collinear (\(r > 0.8\)), applying Principal Component
Analysis (PCA) or pruning redundant features helps tree-naive algorithms
like linear discriminant analysis.| Predictor | Original_Skew | Transformed_Skew | |
|---|---|---|---|
| RI | RI | 1.603 | 1.566 |
| Na | Na | 0.448 | 0.034 |
| Mg | Mg | -1.136 | -1.136 |
| Al | Al | 0.895 | 0.091 |
| Si | Si | -0.720 | -0.651 |
| K | K | 6.460 | 6.460 |
| Ca | Ca | 2.018 | -0.194 |
| Ba | Ba | 3.369 | 3.369 |
| Fe | Fe | 1.730 | 1.730 |
The soybean data can also be found at the UC Irvine Machine Learning Repository. Data were collected to predict disease in 683 soybeans. The 35 predictors are mostly categorical and include information on the environmental conditions (e.g., temperature, precipitation) and plant conditions (e.g., left spots, mold growth). The outcome labels consist of 19 distinct classes.
Note that the echo = FALSE parameter was added to the
code chunk to prevent printing of the R code that generated the
plot.
The Soybean dataset contains 683 samples and 35
categorical predictors (soil, weather, leaf, and plant symptoms)
categorized into 19 disease classes.
A predictor is degenerate (or near-zero variance) if it contains only one unique value, or if the ratio of the most frequent value to the second most frequent value is exceptionally large (\(> 95/5\) rule) and the percentage of unique values is small.
| Predictor | freqRatio | percentUnique | zeroVar | nzv |
|---|---|---|---|---|
| leaf.mild | 26.75 | 0.4392387 | FALSE | TRUE |
| mycelium | 106.50 | 0.2928258 | FALSE | TRUE |
| sclerotia | 31.25 | 0.2928258 | FALSE | TRUE |
leaf.mild, mycelium, and
sclerotia exhibit extreme frequency imbalances (frequency
ratios \(> 15\) to over \(100\)) and almost zero variance.Approximately 18% of the data values are missing across the observations. We evaluate missingness rates per predictor and assess if missingness correlates with specific disease classes.
| Predictor | Percent_Missing |
|---|---|
| hail | 17.71596 |
| sever | 17.71596 |
| seed.tmt | 17.71596 |
| lodging | 17.71596 |
| germ | 16.39824 |
| leaf.mild | 15.81259 |
| fruiting.bodies | 15.51977 |
| fruit.spots | 15.51977 |
| seed.discolor | 15.51977 |
| shriveling | 15.51977 |
| Class | Total_Cases | Samples_With_Missing | Pct_Cases_Incomplete | Mean_Missing_Fields |
|---|---|---|---|---|
| 2-4-d-injury | 16 | 16 | 100.0 | 28.1 |
| cyst-nematode | 14 | 14 | 100.0 | 24.0 |
| diaporthe-pod-&-stem-blight | 15 | 15 | 100.0 | 11.8 |
| herbicide-injury | 8 | 8 | 100.0 | 20.0 |
| phytophthora-rot | 88 | 68 | 77.3 | 13.8 |
| alternarialeaf-spot | 91 | 0 | 0.0 | 0.0 |
| anthracnose | 44 | 0 | 0.0 | 0.0 |
| bacterial-blight | 20 | 0 | 0.0 | 0.0 |
| bacterial-pustule | 20 | 0 | 0.0 | 0.0 |
| brown-spot | 92 | 0 | 0.0 | 0.0 |
| brown-stem-rot | 44 | 0 | 0.0 | 0.0 |
| charcoal-rot | 20 | 0 | 0.0 | 0.0 |
| diaporthe-stem-canker | 20 | 0 | 0.0 | 0.0 |
| downy-mildew | 20 | 0 | 0.0 | 0.0 |
| frog-eye-leaf-spot | 91 | 0 | 0.0 | 0.0 |
| phyllosticta-leaf-spot | 20 | 0 | 0.0 | 0.0 |
| powdery-mildew | 20 | 0 | 0.0 | 0.0 |
| purple-seed-stain | 20 | 0 | 0.0 | 0.0 |
| rhizoctonia-root-rot | 20 | 0 | 0.0 | 0.0 |
2-4-d-injury,
cyst-nematode, diaporthe-stem-canker, and
phytophthora-rot consist of 100% incomplete
observations.bacterial-blight,
brown-spot, and anthracnose contain zero
missing values across all measurements.Do Not Indiscriminately Drop Incomplete Rows:
Dropping all cases with missing data (na.omit()) would
completely eliminate entire diagnostic classes (e.g., eliminating all
occurrences of 2-4-d-injury).
Filter Degenerate Predictors First: Remove
uninformative predictors with near-zero variance (mycelium,
sclerotia, leaf.mild) identified in part
(a).
Imputation Methods:
The official state motto of Texas is “Friendship”.
What it Refers To
Origin of the State’s Name: The motto refers to the origin of the word “Texas” (or the Spanish pronunciation, “Tejas”).
Native American Roots: It comes from a Caddo (specifically Hasinai) Native American word—such as teyshas or texias—which translates to “friends” or “allies”.
Adoption: The 41st Texas Legislature officially adopted “Friendship” as the state motto in February 1930 to honor this historic linguistic connection.