Coffee
Predicting Coffee Rating
Data
Motivation
The motivation for this project is to build a machine learning model that can predict the rating of a coffee based on the properties or characteristics of the coffee bean. Additionally, this project serves as the final project for CS 670(Data Science) course. The dataset comes from Kraggle which was gather in January, 2018 from the Coffee Quality Institute (CQI).
Composition
The data has been preprocessed to removed features that may not be relevant. After processing, the columns/features in the dataset are as follows:
| Variables | Description |
|---|---|
| Species | The species or type that the coffee bean belongs to (e.g., Arabica, Robusta). |
| Country.of.Origin | Country where the coffee bean comes from. |
| Region | Specific region within the country where the coffee is grown. |
| Variety | The specific variety or cultivar of the coffee plant. |
| Processing.Method | The method used to process the coffee beans after harvesting (e.g., washed, natural, honey). |
| Aroma | The fragrance or smell of the coffee, scored during cupping. |
| Flavor | The taste of the coffee, scored during cupping. |
| Aftertaste | The lingering flavor after the coffee is swallowed, scored during cupping. |
| Acidity | The brightness or sharpness of the coffee’s flavor, scored during cupping. |
| Body | The mouthfeel or weight of the coffee, scored during cupping. |
| Balance | The harmony of flavors, acidity, and body in the coffee, scored during cupping. |
| Uniformity | Consistency of the coffee’s flavors across different cups, scored during cupping. |
| Clean.Cup | The absence of defects and clarity of flavor, scored during cupping. |
| Sweetness | The natural sweetness of the coffee, scored during cupping. |
| Cupper.Points | The individual score given by the cupper (taster) based on their overall impression of the coffee. |
| Total.Cup.Points | The total score of the coffee, summing up all the individual attribute scores. |
| Moisture | The moisture content of the coffee beans. |
| Category.One.Defects | The number of primary defects in the coffee sample. |
| Quakers | The number of underdeveloped or defective beans (quakers) in the sample. |
| Category.Two.Defects | The number of secondary defects in the coffee sample. |
| altitude_low_meters | The lower end of the altitude range (in meters) where the coffee is grown. |
| altitude_high_meters | The higher end of the altitude range (in meters) where the coffee is grown. |
Although, this data include sensitive information on coffee producer, it is publicly available and free to use as is not offensive, insulting, or threatening.
### Preprocessing/Cleaning/Labeling
Inside this dataset, there are columns with null values which are altitude_low_meters and altitude_low_meters. For the purpose of consistency, the entries will be assign with the value of 0.
Uses
There are many ways on how to use this data. Below are examples on how this data could be used:
- Predicting the rating of the coffee based on the bean properties.
- Insight on countries that produces “good” coffee.
- Insight on coffee production of several companies.
- Predicting the rating of the coffee based on the altitude.
Distribution
This data is publicly available for use which can be found at https://www.kaggle.com/datasets/volpatto/coffee-quality-database-from-cqi. For a more recent version, this data is available at https://database.coffeeinstitute.org/. However, the latter may require scrapping tools to acquire all the necessary data.
Maintenance
Currently, there is no plan on updating or maintaining the data. The data was gather in January, 2018.
Questions
- Can we predict the rating of a coffee based on other factors that are not flavor related?
- Can we predict the flavor score based on other flavor score?
- Can we predict coffee rating based on the geographic location?
Visualization
| altitude_low_meters | altitude_high_meters | |
|---|---|---|
| Min. : 0.0 | Min. : 0 | |
| 1st Qu.: 760.5 | 1st Qu.: 800 | |
| Median : 1219.2 | Median : 1250 | |
| Mean : 1450.0 | Mean : 1490 | |
| 3rd Qu.: 1500.0 | 3rd Qu.: 1550 | |
| Max. :190164.0 | Max. :190164 |
Based on the previous data, we found entries that does not make sense. Thus, we will remove them.
| altitude_low_meters | altitude_high_meters | |
|---|---|---|
| Min. : 0 | Min. : 0 | |
| 1st Qu.: 754 | 1st Qu.: 800 | |
| Median :1219 | Median :1250 | |
| Mean :1079 | Mean :1119 | |
| 3rd Qu.:1500 | 3rd Qu.:1550 | |
| Max. :4287 | Max. :5900 |
Min. 1st Qu. Median Mean 3rd Qu. Max.
0.00 81.12 82.50 82.09 83.67 90.58
Country.of.Origin Total.Cup.Points
1 Honduras 0
Country.of.Origin Total.Cup.Points
1 Ethiopia 90.58
Top 3 Countries by Average Total Cup Points:
# A tibble: 3 × 2
Country.of.Origin mean_total_cup_points
<chr> <dbl>
1 Japan 84.7
2 Ethiopia 85.5
3 Papua New Guinea 85.8
Bottom 3 Countries by Average Total Cup Points:
# A tibble: 3 × 2
Country.of.Origin mean_total_cup_points
<chr> <dbl>
1 Haiti 77.2
2 Ivory Coast 79.3
3 Honduras 79.4
34 codes from your data successfully matched countries in the map
0 codes from your data failed to match with a country code in the map
209 codes from the map weren't represented in your data
Findings
- In the distribution by species, we can find that the majority of coffee beans are arabica which could be a strong indication that arabica tends to be a have its profile tested.
- In the distribution by country of origin, we can find that Mexico, Guatemala, and Colombia are the top countries to test their product.
- Based on the low and high distribution, we can assume that coffees with an average of 1200 meters of altitute are the must subject to have their profile tested.
- From the total cup points distribution and average total cup points, we can see that generally the coffee that are being test got a good score on every country.
- Although the average cup point for each country is high, we can see that there is one that got 0 (Honduras) which could indicate an outlier.
- Based on the country heat map, we can see that most of the coffee comes from America which could indicate that America coffee is the most common coffee.
From the correlation graph, I found interesting is that the flavor profile have strong correlation to each other. This could indicate that we could predict the flavor profile score based on other flavors. Another thing worth to mention is that defects tends not to affect too much on the overall score, and that the altitude has no correlation with the score which could make difficult to make a prediction model based on these characteristics.
Models
First, we need to randomly split the data into a 0.7 ratio for training and 0.3 ratio for testing.
Logistic Regression
Then, we must convert the data that we want to predict between 0 or 1 based on the threshold. In this case, we considered the coffee to be flavorful if the score is greater than 8.
Flavor Model with One Feature
Below, we train the model to identify the Flavor score solely based on the aroma.
Call:
glm(formula = Flavor ~ Aroma, family = binomial, data = train)
Coefficients:
Estimate Std. Error z value Pr(>|z|)
(Intercept) -65.3913 6.1572 -10.62 <2e-16 ***
Aroma 8.0756 0.7795 10.36 <2e-16 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
(Dispersion parameter for binomial family taken to be 1)
Null deviance: 517.22 on 933 degrees of freedom
Residual deviance: 289.72 on 932 degrees of freedom
AIC: 293.72
Number of Fisher Scoring iterations: 7
We then test the model by applying a confusion matrix.
Confusion Matrix and Statistics
Reference
Prediction 0 1
0 367 19
1 6 9
Accuracy : 0.9377
95% CI : (0.9093, 0.9593)
No Information Rate : 0.9302
P-Value [Acc > NIR] : 0.3198
Kappa : 0.3888
Mcnemar's Test P-Value : 0.0164
Sensitivity : 0.9839
Specificity : 0.3214
Pos Pred Value : 0.9508
Neg Pred Value : 0.6000
Prevalence : 0.9302
Detection Rate : 0.9152
Detection Prevalence : 0.9626
Balanced Accuracy : 0.6527
'Positive' Class : 0
Based on the previous, the model has a high overall accuracy (93.77%) and sensitivity (98.39%), meaning it correctly identifies most of the true positive cases. However, it has low specificity (32.14%), indicating it struggles to correctly identify negative cases.The positive predictive value is high (95.08%), suggesting most of the positive predictions are correct. The balanced accuracy (65.27%) reflects a more balanced view of the model’s performance considering both sensitivity and specificity. Kappa value (0.3888) indicates moderate agreement between the predictions and actual classifications beyond chance. Overall, the model performs well in identifying the positive class but has room for improvement in correctly identifying the negative class.
Flavor Model with Multiple Features
Call:
glm(formula = Flavor ~ Aroma + Aftertaste + Acidity + Body +
Balance + Uniformity + Clean.Cup + Sweetness + Cupper.Points,
family = binomial, data = train)
Coefficients:
Estimate Std. Error z value Pr(>|z|)
(Intercept) -113.07588 13.83069 -8.176 2.94e-16 ***
Aroma 4.62549 0.90154 5.131 2.89e-07 ***
Aftertaste 3.38695 1.23296 2.747 0.00601 **
Acidity 0.62880 0.87230 0.721 0.47100
Body 2.54711 1.13540 2.243 0.02487 *
Balance 0.66003 1.10691 0.596 0.55099
Uniformity -0.12329 0.39665 -0.311 0.75593
Clean.Cup 0.29078 0.48801 0.596 0.55127
Sweetness 0.02909 0.35472 0.082 0.93463
Cupper.Points 2.16559 0.81765 2.649 0.00808 **
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
(Dispersion parameter for binomial family taken to be 1)
Null deviance: 517.22 on 933 degrees of freedom
Residual deviance: 210.34 on 924 degrees of freedom
AIC: 230.34
Number of Fisher Scoring iterations: 8
Confusion Matrix and Statistics
Reference
Prediction 0 1
0 370 11
1 3 17
Accuracy : 0.9651
95% CI : (0.9421, 0.9808)
No Information Rate : 0.9302
P-Value [Acc > NIR] : 0.002076
Kappa : 0.6903
Mcnemar's Test P-Value : 0.061369
Sensitivity : 0.9920
Specificity : 0.6071
Pos Pred Value : 0.9711
Neg Pred Value : 0.8500
Prevalence : 0.9302
Detection Rate : 0.9227
Detection Prevalence : 0.9501
Balanced Accuracy : 0.7995
'Positive' Class : 0
Based on the previous confusion matrix, the model shows significant improvement across all key metrics after adding more features. Accuracy, kappa, sensitivity, specificity, positive and negative predictive values, and balanced accuracy have all increased, indicating that the model is better at both identifying true positives and true negatives, and overall classification performance has improved.
Flavor Model with External Features
Now, we will try to predict the flavor score with other external features such as geolocation, altitude, and/or processing methods.From the previous, correlation matrix, we found that latitudes have little or no correlation with other features. However, we only tested for linear correlation. Thus, we may need to explore more in details about the data.
From the previous, we observe that Aroma and Flavor are generally higher at mid to high altitudes, suggesting that these quality attributes improve with higher altitude. Body and Acidity might show a wide spread at all altitudes, indicating that these attributes are more variable and might be influenced by factors other than altitude.Uniformity and Clean Cup show tight clustering at all altitudes, suggesting high consistency regardless of altitude. Thus, Aroma and Flavor seems to have a some relation with altitude.
Call:
glm(formula = Flavor ~ altitude_low_meters, family = binomial,
data = train)
Coefficients:
Estimate Std. Error z value Pr(>|z|)
(Intercept) -2.7671673 0.2500613 -11.066 <2e-16 ***
altitude_low_meters 0.0002810 0.0001872 1.501 0.133
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
(Dispersion parameter for binomial family taken to be 1)
Null deviance: 517.22 on 933 degrees of freedom
Residual deviance: 514.95 on 932 degrees of freedom
AIC: 518.95
Number of Fisher Scoring iterations: 5
Confusion Matrix and Statistics
Reference
Prediction 0 1
0 373 28
1 0 0
Accuracy : 0.9302
95% CI : (0.9007, 0.9531)
No Information Rate : 0.9302
P-Value [Acc > NIR] : 0.5501
Kappa : 0
Mcnemar's Test P-Value : 3.352e-07
Sensitivity : 1.0000
Specificity : 0.0000
Pos Pred Value : 0.9302
Neg Pred Value : NaN
Prevalence : 0.9302
Detection Rate : 0.9302
Detection Prevalence : 1.0000
Balanced Accuracy : 0.5000
'Positive' Class : 0
From the previous, we can see that the logistic regression model is heavily biased towards predicting the majority class (class 0), with no instances of the minority class (class 1) being correctly predicted. This indicates that this model is no good to predict flavor score using altitude.
Random Forest
It seems that the flavor cannot be predicted using altitute in a linear manner. Thus, let’s try to use models that are not linear. In this case random forest.
Confusion Matrix and Statistics
Reference
Prediction 0 6.08 6.17 6.33 6.42 6.5 6.58 6.67 6.75 6.83 6.92 7 7.08 7.17
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.75 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.83 0 0 0 0 0 0 0 0 0 0 0 1 0 0
6.92 0 0 0 0 0 0 0 0 0 0 0 0 0 1
7 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.25 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.33 0 0 0 0 0 1 0 0 0 0 1 0 0 2
7.42 0 0 0 0 0 0 0 0 1 0 0 0 0 1
7.5 0 0 0 0 0 0 1 0 0 1 2 2 2 4
7.58 0 0 0 0 0 0 0 1 1 2 0 3 3 2
7.67 0 0 0 0 0 0 0 0 0 0 0 1 0 0
7.75 0 0 0 0 0 0 0 0 0 0 0 0 1 0
7.81 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.83 0 0 0 0 0 0 0 0 0 0 0 0 0 1
7.88 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.92 0 0 0 0 0 0 0 0 0 0 0 0 2 0
8 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.25 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.83 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Reference
Prediction 7.25 7.33 7.42 7.5 7.58 7.67 7.75 7.81 7.83 7.88 7.92 8 8.08 8.17
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.58 0 0 0 0 0 1 0 0 0 0 0 0 0 0
6.67 0 0 0 0 0 0 0 0 0 0 0 1 0 0
6.75 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.83 0 0 0 0 0 0 0 0 0 0 1 0 0 0
6.92 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.08 0 1 0 1 1 0 1 0 0 0 0 0 0 0
7.17 0 1 1 0 0 0 0 0 0 0 0 0 0 0
7.25 0 0 0 0 1 0 0 0 0 0 0 0 0 0
7.33 1 1 1 4 5 4 2 0 1 0 1 0 0 0
7.42 0 1 3 4 2 2 1 0 2 0 2 1 0 0
7.5 7 7 8 10 6 1 7 0 3 0 1 2 1 0
7.58 3 7 5 9 12 10 12 0 9 0 4 3 0 3
7.67 0 2 1 1 3 6 1 0 2 0 0 1 0 0
7.75 1 0 2 3 3 5 0 0 0 0 0 0 0 0
7.81 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.83 0 1 1 0 0 0 0 0 0 0 0 0 0 0
7.88 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.92 0 0 0 1 0 0 1 0 0 0 0 0 1 0
8 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.25 0 1 0 0 0 0 0 0 0 0 0 0 0 0
8.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.83 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Reference
Prediction 8.25 8.33 8.42 8.5 8.58 8.67 8.83
0 0 0 0 0 0 0 0
6.08 0 0 0 0 0 0 0
6.17 0 0 0 0 0 0 0
6.33 0 0 0 0 0 0 0
6.42 0 0 0 0 0 0 0
6.5 0 0 0 0 0 0 0
6.58 0 0 0 0 0 0 0
6.67 0 0 0 0 0 0 0
6.75 0 0 0 0 0 0 0
6.83 0 0 0 0 0 0 0
6.92 0 0 0 0 0 0 0
7 0 0 0 0 0 0 0
7.08 0 0 0 0 0 0 0
7.17 0 0 0 0 0 0 0
7.25 0 0 0 0 0 0 0
7.33 0 0 1 0 0 0 0
7.42 0 0 0 0 0 0 0
7.5 0 0 0 0 0 0 0
7.58 1 0 0 1 0 0 0
7.67 0 0 0 0 0 0 0
7.75 0 0 0 0 0 0 0
7.81 0 0 0 0 0 0 0
7.83 0 0 0 0 0 0 0
7.88 0 0 0 0 0 0 0
7.92 0 0 0 0 0 0 0
8 0 0 0 0 0 0 0
8.08 0 0 0 0 0 0 0
8.17 0 0 0 0 0 0 0
8.25 0 0 0 0 0 0 0
8.33 0 1 0 0 0 0 0
8.42 0 0 0 0 0 0 0
8.5 0 0 0 0 0 0 0
8.58 0 0 0 0 0 0 0
8.67 0 0 0 0 0 0 0
8.83 0 0 0 0 0 0 0
Overall Statistics
Accuracy : 0.1289
95% CI : (0.0904, 0.1762)
No Information Rate : 0.1289
P-Value [Acc > NIR] : 0.528
Kappa : 0.0212
Mcnemar's Test P-Value : NA
Statistics by Class:
Class: 0 Class: 6.08 Class: 6.17 Class: 6.33 Class: 6.42
Sensitivity NA NA NA NA NA
Specificity 1 1 1 1 1
Pos Pred Value NA NA NA NA NA
Neg Pred Value NA NA NA NA NA
Prevalence 0 0 0 0 0
Detection Rate 0 0 0 0 0
Detection Prevalence 0 0 0 0 0
Balanced Accuracy NA NA NA NA NA
Class: 6.5 Class: 6.58 Class: 6.67 Class: 6.75 Class: 6.83
Sensitivity 0.000000 0.000000 0.000000 0.000000 0.000000
Specificity 1.000000 0.996078 0.996078 1.000000 0.992095
Pos Pred Value NaN 0.000000 0.000000 NaN 0.000000
Neg Pred Value 0.996094 0.996078 0.996078 0.992188 0.988189
Prevalence 0.003906 0.003906 0.003906 0.007812 0.011719
Detection Rate 0.000000 0.000000 0.000000 0.000000 0.000000
Detection Prevalence 0.000000 0.003906 0.003906 0.000000 0.007812
Balanced Accuracy 0.500000 0.498039 0.498039 0.500000 0.496047
Class: 6.92 Class: 7 Class: 7.08 Class: 7.17 Class: 7.25
Sensitivity 0.000000 0.00000 0.00000 0.000000 0.000000
Specificity 0.996047 1.00000 0.98387 0.991837 0.995902
Pos Pred Value 0.000000 NaN 0.00000 0.000000 0.000000
Neg Pred Value 0.988235 0.97266 0.96825 0.956693 0.952941
Prevalence 0.011719 0.02734 0.03125 0.042969 0.046875
Detection Rate 0.000000 0.00000 0.00000 0.000000 0.000000
Detection Prevalence 0.003906 0.00000 0.01562 0.007812 0.003906
Balanced Accuracy 0.498024 0.50000 0.49194 0.495918 0.497951
Class: 7.33 Class: 7.42 Class: 7.5 Class: 7.58 Class: 7.67
Sensitivity 0.045455 0.13636 0.30303 0.36364 0.20690
Specificity 0.897436 0.92735 0.75336 0.64574 0.94714
Pos Pred Value 0.040000 0.15000 0.15385 0.13187 0.33333
Neg Pred Value 0.909091 0.91949 0.87958 0.87273 0.90336
Prevalence 0.085938 0.08594 0.12891 0.12891 0.11328
Detection Rate 0.003906 0.01172 0.03906 0.04688 0.02344
Detection Prevalence 0.097656 0.07812 0.25391 0.35547 0.07031
Balanced Accuracy 0.471445 0.53186 0.52820 0.50469 0.57702
Class: 7.75 Class: 7.81 Class: 7.83 Class: 7.88
Sensitivity 0.00000 NA 0.00000 NA
Specificity 0.93506 1 0.98745 1
Pos Pred Value 0.00000 NA 0.00000 NA
Neg Pred Value 0.89627 NA 0.93281 NA
Prevalence 0.09766 0 0.06641 0
Detection Rate 0.00000 0 0.00000 0
Detection Prevalence 0.05859 0 0.01172 0
Balanced Accuracy 0.46753 NA 0.49372 NA
Class: 7.92 Class: 8 Class: 8.08 Class: 8.17 Class: 8.25
Sensitivity 0.00000 0.00000 0.000000 0.00000 0.000000
Specificity 0.97976 1.00000 1.000000 1.00000 0.996078
Pos Pred Value 0.00000 NaN NaN NaN 0.000000
Neg Pred Value 0.96414 0.96875 0.992188 0.98828 0.996078
Prevalence 0.03516 0.03125 0.007812 0.01172 0.003906
Detection Rate 0.00000 0.00000 0.000000 0.00000 0.000000
Detection Prevalence 0.01953 0.00000 0.000000 0.00000 0.003906
Balanced Accuracy 0.48988 0.50000 0.500000 0.50000 0.498039
Class: 8.33 Class: 8.42 Class: 8.5 Class: 8.58 Class: 8.67
Sensitivity 1.000000 0.000000 0.000000 NA NA
Specificity 1.000000 1.000000 1.000000 1 1
Pos Pred Value 1.000000 NaN NaN NA NA
Neg Pred Value 1.000000 0.996094 0.996094 NA NA
Prevalence 0.003906 0.003906 0.003906 0 0
Detection Rate 0.003906 0.000000 0.000000 0 0
Detection Prevalence 0.003906 0.000000 0.000000 0 0
Balanced Accuracy 1.000000 0.500000 0.500000 NA NA
Class: 8.83
Sensitivity NA
Specificity 1
Pos Pred Value NA
Neg Pred Value NA
Prevalence 0
Detection Rate 0
Detection Prevalence 0
Balanced Accuracy NA
Based on the previous, we managed to get rid of the bias from the previous model. However, the accuracy is too low. Hence, not making a very good model to predict flavor, at least just based on altitudes.
Random Forest with more features
Lets try to add more features that are not flavor related such as moisture and defects.
Confusion Matrix and Statistics
Reference
Prediction 0 6.08 6.17 6.33 6.42 6.5 6.58 6.67 6.75 6.83 6.92 7 7.08 7.17 7.25
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.75 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.83 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1
6.92 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0
7 0 0 0 0 0 0 0 0 1 0 0 1 2 0 1
7.08 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0
7.17 0 0 0 0 0 0 0 0 0 1 1 0 0 1 0
7.25 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0
7.33 0 0 0 0 0 0 1 0 0 1 1 2 1 0 2
7.42 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0
7.5 0 0 0 0 0 0 0 0 0 0 1 1 0 2 4
7.58 0 0 0 0 0 0 0 0 0 1 0 1 3 0 0
7.67 0 0 0 0 0 0 0 0 0 0 0 0 0 2 2
7.75 0 0 0 0 0 0 0 1 0 0 0 1 0 1 1
7.81 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.83 0 0 0 0 0 0 0 0 1 0 0 0 0 1 0
7.88 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.92 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0
8 0 0 0 0 0 1 0 0 0 0 0 1 0 0 1
8.08 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0
8.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.25 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.83 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Reference
Prediction 7.33 7.42 7.5 7.58 7.67 7.75 7.81 7.83 7.88 7.92 8 8.08 8.17 8.25
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.75 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.83 0 0 0 0 0 0 0 0 0 0 0 0 0 0
6.92 0 0 0 0 0 1 0 0 0 0 0 0 0 0
7 1 0 0 0 0 0 0 0 0 0 0 0 0 0
7.08 1 1 0 1 1 0 0 1 0 0 0 0 0 0
7.17 1 1 0 2 0 2 0 0 0 1 0 0 0 0
7.25 0 0 2 3 0 3 0 1 0 1 0 0 0 0
7.33 0 1 3 3 5 0 0 1 0 0 0 0 0 0
7.42 2 1 1 1 1 1 0 1 0 2 2 0 0 1
7.5 7 4 4 7 4 6 0 4 0 2 1 0 1 0
7.58 7 5 8 6 8 5 0 4 0 0 2 1 1 0
7.67 3 4 6 5 2 4 0 0 0 1 1 0 0 0
7.75 0 4 8 4 3 3 0 3 0 1 1 0 0 0
7.81 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.83 0 1 0 1 2 0 0 2 0 1 1 1 0 0
7.88 0 0 0 0 0 0 0 0 0 0 0 0 0 0
7.92 0 0 1 0 1 0 0 0 0 0 0 0 0 0
8 0 0 0 0 2 0 0 0 0 0 0 0 1 0
8.08 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.17 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.25 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.33 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.42 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.5 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.58 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.67 0 0 0 0 0 0 0 0 0 0 0 0 0 0
8.83 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Reference
Prediction 8.33 8.42 8.5 8.58 8.67 8.83
0 0 0 0 0 0 0
6.08 0 0 0 0 0 0
6.17 0 0 0 0 0 0
6.33 0 0 0 0 0 0
6.42 0 0 0 0 0 0
6.5 0 0 0 0 0 0
6.58 0 0 0 0 0 0
6.67 0 0 0 0 0 0
6.75 0 0 0 0 0 0
6.83 0 0 0 0 0 0
6.92 0 0 0 0 0 0
7 0 0 0 0 0 0
7.08 0 0 0 0 0 0
7.17 0 0 0 0 0 0
7.25 0 0 0 0 0 0
7.33 0 0 0 0 0 0
7.42 0 0 0 0 0 0
7.5 0 0 0 0 0 0
7.58 0 0 0 0 0 0
7.67 1 0 0 0 0 0
7.75 0 1 0 0 0 0
7.81 0 0 0 0 0 0
7.83 0 0 0 0 0 0
7.88 0 0 0 0 0 0
7.92 0 0 0 0 0 0
8 0 0 0 0 0 0
8.08 0 0 0 0 0 0
8.17 0 0 0 0 0 0
8.25 0 0 0 0 0 0
8.33 0 0 0 0 0 0
8.42 0 0 0 0 0 0
8.5 0 0 0 0 0 0
8.58 0 0 0 0 0 0
8.67 0 0 1 0 0 0
8.83 0 0 0 0 0 0
Overall Statistics
Accuracy : 0.082
95% CI : (0.0515, 0.1227)
No Information Rate : 0.1289
P-Value [Acc > NIR] : 0.993
Kappa : -0.0169
Mcnemar's Test P-Value : NA
Statistics by Class:
Class: 0 Class: 6.08 Class: 6.17 Class: 6.33 Class: 6.42
Sensitivity NA NA NA NA NA
Specificity 1 1 1 1 1
Pos Pred Value NA NA NA NA NA
Neg Pred Value NA NA NA NA NA
Prevalence 0 0 0 0 0
Detection Rate 0 0 0 0 0
Detection Prevalence 0 0 0 0 0
Balanced Accuracy NA NA NA NA NA
Class: 6.5 Class: 6.58 Class: 6.67 Class: 6.75 Class: 6.83
Sensitivity 0.000000 0.000000 0.000000 0.000000 0.000000
Specificity 1.000000 1.000000 1.000000 1.000000 0.996047
Pos Pred Value NaN NaN NaN NaN 0.000000
Neg Pred Value 0.996094 0.996094 0.996094 0.992188 0.988235
Prevalence 0.003906 0.003906 0.003906 0.007812 0.011719
Detection Rate 0.000000 0.000000 0.000000 0.000000 0.000000
Detection Prevalence 0.000000 0.000000 0.000000 0.000000 0.003906
Balanced Accuracy 0.500000 0.500000 0.500000 0.500000 0.498024
Class: 6.92 Class: 7 Class: 7.08 Class: 7.17 Class: 7.25
Sensitivity 0.000000 0.142857 0.125000 0.090909 0.00000
Specificity 0.992095 0.979920 0.979839 0.963265 0.95492
Pos Pred Value 0.000000 0.166667 0.166667 0.100000 0.00000
Neg Pred Value 0.988189 0.976000 0.972000 0.959350 0.95102
Prevalence 0.011719 0.027344 0.031250 0.042969 0.04688
Detection Rate 0.000000 0.003906 0.003906 0.003906 0.00000
Detection Prevalence 0.007812 0.023438 0.023438 0.039062 0.04297
Balanced Accuracy 0.496047 0.561388 0.552419 0.527087 0.47746
Class: 7.33 Class: 7.42 Class: 7.5 Class: 7.58 Class: 7.67
Sensitivity 0.00000 0.045455 0.12121 0.18182 0.068966
Specificity 0.91026 0.944444 0.80269 0.79372 0.872247
Pos Pred Value 0.00000 0.071429 0.08333 0.11538 0.064516
Neg Pred Value 0.90638 0.913223 0.86058 0.86765 0.880000
Prevalence 0.08594 0.085938 0.12891 0.12891 0.113281
Detection Rate 0.00000 0.003906 0.01562 0.02344 0.007812
Detection Prevalence 0.08203 0.054688 0.18750 0.20312 0.121094
Balanced Accuracy 0.45513 0.494949 0.46195 0.48777 0.470606
Class: 7.75 Class: 7.81 Class: 7.83 Class: 7.88
Sensitivity 0.12000 NA 0.117647 NA
Specificity 0.87446 1 0.962343 1
Pos Pred Value 0.09375 NA 0.181818 NA
Neg Pred Value 0.90179 NA 0.938776 NA
Prevalence 0.09766 0 0.066406 0
Detection Rate 0.01172 0 0.007812 0
Detection Prevalence 0.12500 0 0.042969 0
Balanced Accuracy 0.49723 NA 0.539995 NA
Class: 7.92 Class: 8 Class: 8.08 Class: 8.17 Class: 8.25
Sensitivity 0.00000 0.00000 0.000000 0.00000 0.000000
Specificity 0.98785 0.97581 0.996063 1.00000 1.000000
Pos Pred Value 0.00000 0.00000 0.000000 NaN NaN
Neg Pred Value 0.96443 0.96800 0.992157 0.98828 0.996094
Prevalence 0.03516 0.03125 0.007812 0.01172 0.003906
Detection Rate 0.00000 0.00000 0.000000 0.00000 0.000000
Detection Prevalence 0.01172 0.02344 0.003906 0.00000 0.000000
Balanced Accuracy 0.49393 0.48790 0.498031 0.50000 0.500000
Class: 8.33 Class: 8.42 Class: 8.5 Class: 8.58 Class: 8.67
Sensitivity 0.000000 0.000000 0.000000 NA NA
Specificity 1.000000 1.000000 1.000000 1 0.996094
Pos Pred Value NaN NaN NaN NA NA
Neg Pred Value 0.996094 0.996094 0.996094 NA NA
Prevalence 0.003906 0.003906 0.003906 0 0.000000
Detection Rate 0.000000 0.000000 0.000000 0 0.000000
Detection Prevalence 0.000000 0.000000 0.000000 0 0.003906
Balanced Accuracy 0.500000 0.500000 0.500000 NA NA
Class: 8.83
Sensitivity NA
Specificity 1
Pos Pred Value NA
Neg Pred Value NA
Prevalence 0
Detection Rate 0
Detection Prevalence 0
Balanced Accuracy NA
By adding more features, we actually did worse. Because of this, we can assume that there is no enough data to make a good use of random forest.
Result, Analysis, and Discussion
The logistic regression models demonstrated varying degrees of effectiveness in predicting coffee flavor scores based on different sets of features:
Results
Single Feature (Aroma):
Accuracy: 93.53%
Sensitivity: 98.13%
Specificity: 32.14%
Positive Predictive Value: 95.08%
Kappa: 0.3776
Multiple Features (Aroma, Aftertaste, Acidity, Body, Balance, Uniformity, Clean Cup, Sweetness):
Accuracy: 96.27%
Sensitivity: 98.93%
Specifity: 60.71%
Positive Predictive Value: 97.11%
Kappa: 0.6744
(Improved metrics across the board, indicating a more robust model)
External Features (Altitude):
The model failed to predict the minority class (Flavor score of 1), highlighting the inadequacy of altitude as a predictive feature for coffee flavor.
Random Forest
The model accuracy is incredibly low. This indicates that the model is under performing. Thus, in order to improve the model more data is needed.
Analysis
The initial exploratory data analysis (EDA) revealed several important insights:
Species Distribution: The majority of coffee samples are Arabica, suggesting a focus on testing this species.
Country of Origin Distribution: Mexico, Guatemala, and Colombia are the top contributors to the dataset, indicating a concentration of coffee quality testing in these regions.
Altitude Distribution: Most coffee samples are grown at an average altitude of 1200 meters, aligning with known optimal conditions for high-quality coffee.
Total Cup Points Distribution: The overall high scores across countries suggest a generally high quality of tested coffee samples, with notable outliers like Honduras with a score of 0, indicating potential data anomalies.
Correlation Matrix: Strong correlations were observed among flavor profile attributes, indicating the possibility of predicting one flavor score based on others. Defects and altitude showed little to no correlation with flavor scores, suggesting limited predictive power for these features.
Linear Models: Linear model seems to perform very well for features with strong correlation in this case predicting flavor score based on other flavor profiles.
Non-linear Models: Random forest did not perform well when trying to predict the flavor score based on features with low correlation in this case non-flavor based features (altitudes, moisture, defects). These could indicate that we lack of data in order to accurately predict flavor score based on non-flavor information.
Discussion
Predictive Modeling:
The logistic regression model using Aroma as the sole predictor of Flavor achieved high accuracy and sensitivity but struggled with specificity. This indicates that while Aroma is a strong indicator of coffee flavor, it may not be sufficient on its own for comprehensive prediction.
Adding multiple flavor-related features significantly improved the model’s performance, highlighting the importance of a multifaceted approach when predicting complex attributes like coffee flavor.
Using external features such as altitude for flavor prediction proved ineffective. This aligns with the observed lack of correlation between altitude and flavor scores, suggesting that altitude alone does not capture the nuances influencing coffee flavor.
Implications:
The strong inter-correlation among flavor attributes implies that coffee quality assessments are inherently interconnected. Models predicting coffee ratings or specific flavor scores should incorporate multiple related features to enhance accuracy.
The data suggests a geographical bias towards certain coffee-producing countries. This could reflect the industry’s focus on these regions or data collection biases.
The presence of outliers and anomalies, such as the zero scores, indicates a need for further data cleaning and validation to ensure the reliability of predictive models.
Future Work:
Exploring other non-linear models and more complex machine learning techniques may yield better predictions, especially when incorporating diverse features like processing methods and geographical data. Additional data collection from underrepresented regions and varieties could provide a more balanced dataset, enhancing model generalizability. Investigating other external factors, such as soil quality, weather conditions, and farming practices, could offer deeper insights into the determinants of coffee quality.
Impact
This study could have several significant impacts on the coffee industry and Sustainability and Ethical Considerations:
- Coffee Industry:
- Improved Quality Control: The findings from this study can help coffee producers and quality control teams identify the key attributes that contribute to high coffee ratings. By understanding the significance of factors such as aroma, aftertaste, and balance, producers can refine their processes to enhance these characteristics, thereby improving the overall quality of their coffee.
- Enhanced Marketability: Coffee brands can use these insights to market their products more effectively. Highlighting specific flavor attributes and the conditions under which their coffee is grown can create compelling narratives that attract consumers who are looking for high-quality coffee. This can differentiate products in a competitive market and build brand loyalty.
- Geographical Insights: The analysis identifies regions that consistently produce high-quality coffee. This information can guide investment decisions and resource allocation for coffee producers and traders. Promoting regions with a track record of excellence can enhance their reputation and market share, while discovering new high-quality coffee sources can expand the diversity of offerings.
- Sustainability and Ethical Considerations:
Sustainable Practices: By pinpointing the key factors that lead to high-quality coffee, the study can guide sustainable agricultural practices. Coffee producers can optimize these factors to produce high-quality coffee with minimal environmental impact, promoting sustainability within the industry. This approach supports the long-term viability of coffee farming and helps preserve the environment.
Data Transparency and Accessibility: The use of publicly available data sets a precedent for transparency and accessibility in research. By making data accessible, this study encourages other researchers to use and share data, fostering collaboration and innovation in the field. This openness can lead to more comprehensive studies and a deeper understanding of coffee quality, benefiting the entire coffee industry and consumers alike.