Coffee

Author

Alex Xiao Cai

Predicting Coffee Rating

Data

Motivation

The motivation for this project is to build a machine learning model that can predict the rating of a coffee based on the properties or characteristics of the coffee bean. Additionally, this project serves as the final project for CS 670(Data Science) course. The dataset comes from Kraggle which was gather in January, 2018 from the Coffee Quality Institute (CQI).

Composition

The data has been preprocessed to removed features that may not be relevant. After processing, the columns/features in the dataset are as follows:

Variables Description
Species The species or type that the coffee bean belongs to (e.g., Arabica, Robusta).
Country.of.Origin Country where the coffee bean comes from.
Region Specific region within the country where the coffee is grown.
Variety The specific variety or cultivar of the coffee plant.
Processing.Method The method used to process the coffee beans after harvesting (e.g., washed, natural, honey).
Aroma The fragrance or smell of the coffee, scored during cupping.
Flavor The taste of the coffee, scored during cupping.
Aftertaste The lingering flavor after the coffee is swallowed, scored during cupping.
Acidity The brightness or sharpness of the coffee’s flavor, scored during cupping.
Body The mouthfeel or weight of the coffee, scored during cupping.
Balance The harmony of flavors, acidity, and body in the coffee, scored during cupping.
Uniformity Consistency of the coffee’s flavors across different cups, scored during cupping.
Clean.Cup The absence of defects and clarity of flavor, scored during cupping.
Sweetness The natural sweetness of the coffee, scored during cupping.
Cupper.Points The individual score given by the cupper (taster) based on their overall impression of the coffee.
Total.Cup.Points The total score of the coffee, summing up all the individual attribute scores.
Moisture The moisture content of the coffee beans.
Category.One.Defects The number of primary defects in the coffee sample.
Quakers The number of underdeveloped or defective beans (quakers) in the sample.
Category.Two.Defects The number of secondary defects in the coffee sample.
altitude_low_meters The lower end of the altitude range (in meters) where the coffee is grown.
altitude_high_meters The higher end of the altitude range (in meters) where the coffee is grown.

Although, this data include sensitive information on coffee producer, it is publicly available and free to use as is not offensive, insulting, or threatening.

  ### Preprocessing/Cleaning/Labeling

Inside this dataset, there are columns with null values which are altitude_low_meters and altitude_low_meters. For the purpose of consistency, the entries will be assign with the value of 0.

Uses

There are many ways on how to use this data. Below are examples on how this data could be used:

  1. Predicting the rating of the coffee based on the bean properties.
  2. Insight on countries that produces “good” coffee.
  3. Insight on coffee production of several companies.
  4. Predicting the rating of the coffee based on the altitude.

Distribution

This data is publicly available for use which can be found at https://www.kaggle.com/datasets/volpatto/coffee-quality-database-from-cqi. For a more recent version, this data is available at https://database.coffeeinstitute.org/. However, the latter may require scrapping tools to acquire all the necessary data.

Maintenance

Currently, there is no plan on updating or maintaining the data. The data was gather in January, 2018.

Questions

  1. Can we predict the rating of a coffee based on other factors that are not flavor related?
  2. Can we predict the flavor score based on other flavor score?
  3. Can we predict coffee rating based on the geographic location?

Visualization

Summary Statistics (Altitudes)
altitude_low_meters altitude_high_meters
Min. : 0.0 Min. : 0
1st Qu.: 760.5 1st Qu.: 800
Median : 1219.2 Median : 1250
Mean : 1450.0 Mean : 1490
3rd Qu.: 1500.0 3rd Qu.: 1550
Max. :190164.0 Max. :190164

Based on the previous data, we found entries that does not make sense. Thus, we will remove them.

Summary Statistics (Altitudes)
altitude_low_meters altitude_high_meters
Min. : 0 Min. : 0
1st Qu.: 754 1st Qu.: 800
Median :1219 Median :1250
Mean :1079 Mean :1119
3rd Qu.:1500 3rd Qu.:1550
Max. :4287 Max. :5900

   Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
   0.00   81.12   82.50   82.09   83.67   90.58 
  Country.of.Origin Total.Cup.Points
1          Honduras                0
  Country.of.Origin Total.Cup.Points
1          Ethiopia            90.58

Top 3 Countries by Average Total Cup Points:

# A tibble: 3 × 2
  Country.of.Origin mean_total_cup_points
  <chr>                             <dbl>
1 Japan                              84.7
2 Ethiopia                           85.5
3 Papua New Guinea                   85.8

Bottom 3 Countries by Average Total Cup Points:

# A tibble: 3 × 2
  Country.of.Origin mean_total_cup_points
  <chr>                             <dbl>
1 Haiti                              77.2
2 Ivory Coast                        79.3
3 Honduras                           79.4
34 codes from your data successfully matched countries in the map
0 codes from your data failed to match with a country code in the map
209 codes from the map weren't represented in your data

Findings

  1. In the distribution by species, we can find that the majority of coffee beans are arabica which could be a strong indication that arabica tends to be a have its profile tested.
  2. In the distribution by country of origin, we can find that Mexico, Guatemala, and Colombia are the top countries to test their product.
  3. Based on the low and high distribution, we can assume that coffees with an average of 1200 meters of altitute are the must subject to have their profile tested.
  4. From the total cup points distribution and average total cup points, we can see that generally the coffee that are being test got a good score on every country.
  5. Although the average cup point for each country is high, we can see that there is one that got 0 (Honduras) which could indicate an outlier.
  6. Based on the country heat map, we can see that most of the coffee comes from America which could indicate that America coffee is the most common coffee.

From the correlation graph, I found interesting is that the flavor profile have strong correlation to each other. This could indicate that we could predict the flavor profile score based on other flavors. Another thing worth to mention is that defects tends not to affect too much on the overall score, and that the altitude has no correlation with the score which could make difficult to make a prediction model based on these characteristics.

Models

First, we need to randomly split the data into a 0.7 ratio for training and 0.3 ratio for testing.

Logistic Regression

Then, we must convert the data that we want to predict between 0 or 1 based on the threshold. In this case, we considered the coffee to be flavorful if the score is greater than 8.

Flavor Model with One Feature

Below, we train the model to identify the Flavor score solely based on the aroma.


Call:
glm(formula = Flavor ~ Aroma, family = binomial, data = train)

Coefficients:
            Estimate Std. Error z value Pr(>|z|)    
(Intercept) -65.3913     6.1572  -10.62   <2e-16 ***
Aroma         8.0756     0.7795   10.36   <2e-16 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

(Dispersion parameter for binomial family taken to be 1)

    Null deviance: 517.22  on 933  degrees of freedom
Residual deviance: 289.72  on 932  degrees of freedom
AIC: 293.72

Number of Fisher Scoring iterations: 7

We then test the model by applying a confusion matrix.

Confusion Matrix and Statistics

          Reference
Prediction   0   1
         0 367  19
         1   6   9
                                          
               Accuracy : 0.9377          
                 95% CI : (0.9093, 0.9593)
    No Information Rate : 0.9302          
    P-Value [Acc > NIR] : 0.3198          
                                          
                  Kappa : 0.3888          
                                          
 Mcnemar's Test P-Value : 0.0164          
                                          
            Sensitivity : 0.9839          
            Specificity : 0.3214          
         Pos Pred Value : 0.9508          
         Neg Pred Value : 0.6000          
             Prevalence : 0.9302          
         Detection Rate : 0.9152          
   Detection Prevalence : 0.9626          
      Balanced Accuracy : 0.6527          
                                          
       'Positive' Class : 0               
                                          

Based on the previous, the model has a high overall accuracy (93.77%) and sensitivity (98.39%), meaning it correctly identifies most of the true positive cases. However, it has low specificity (32.14%), indicating it struggles to correctly identify negative cases.The positive predictive value is high (95.08%), suggesting most of the positive predictions are correct. The balanced accuracy (65.27%) reflects a more balanced view of the model’s performance considering both sensitivity and specificity. Kappa value (0.3888) indicates moderate agreement between the predictions and actual classifications beyond chance. Overall, the model performs well in identifying the positive class but has room for improvement in correctly identifying the negative class.

Flavor Model with Multiple Features


Call:
glm(formula = Flavor ~ Aroma + Aftertaste + Acidity + Body + 
    Balance + Uniformity + Clean.Cup + Sweetness + Cupper.Points, 
    family = binomial, data = train)

Coefficients:
                Estimate Std. Error z value Pr(>|z|)    
(Intercept)   -113.07588   13.83069  -8.176 2.94e-16 ***
Aroma            4.62549    0.90154   5.131 2.89e-07 ***
Aftertaste       3.38695    1.23296   2.747  0.00601 ** 
Acidity          0.62880    0.87230   0.721  0.47100    
Body             2.54711    1.13540   2.243  0.02487 *  
Balance          0.66003    1.10691   0.596  0.55099    
Uniformity      -0.12329    0.39665  -0.311  0.75593    
Clean.Cup        0.29078    0.48801   0.596  0.55127    
Sweetness        0.02909    0.35472   0.082  0.93463    
Cupper.Points    2.16559    0.81765   2.649  0.00808 ** 
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

(Dispersion parameter for binomial family taken to be 1)

    Null deviance: 517.22  on 933  degrees of freedom
Residual deviance: 210.34  on 924  degrees of freedom
AIC: 230.34

Number of Fisher Scoring iterations: 8
Confusion Matrix and Statistics

          Reference
Prediction   0   1
         0 370  11
         1   3  17
                                          
               Accuracy : 0.9651          
                 95% CI : (0.9421, 0.9808)
    No Information Rate : 0.9302          
    P-Value [Acc > NIR] : 0.002076        
                                          
                  Kappa : 0.6903          
                                          
 Mcnemar's Test P-Value : 0.061369        
                                          
            Sensitivity : 0.9920          
            Specificity : 0.6071          
         Pos Pred Value : 0.9711          
         Neg Pred Value : 0.8500          
             Prevalence : 0.9302          
         Detection Rate : 0.9227          
   Detection Prevalence : 0.9501          
      Balanced Accuracy : 0.7995          
                                          
       'Positive' Class : 0               
                                          

Based on the previous confusion matrix, the model shows significant improvement across all key metrics after adding more features. Accuracy, kappa, sensitivity, specificity, positive and negative predictive values, and balanced accuracy have all increased, indicating that the model is better at both identifying true positives and true negatives, and overall classification performance has improved.

Flavor Model with External Features

Now, we will try to predict the flavor score with other external features such as geolocation, altitude, and/or processing methods.From the previous, correlation matrix, we found that latitudes have little or no correlation with other features. However, we only tested for linear correlation. Thus, we may need to explore more in details about the data.

From the previous, we observe that Aroma and Flavor are generally higher at mid to high altitudes, suggesting that these quality attributes improve with higher altitude. Body and Acidity might show a wide spread at all altitudes, indicating that these attributes are more variable and might be influenced by factors other than altitude.Uniformity and Clean Cup show tight clustering at all altitudes, suggesting high consistency regardless of altitude. Thus, Aroma and Flavor seems to have a some relation with altitude.


Call:
glm(formula = Flavor ~ altitude_low_meters, family = binomial, 
    data = train)

Coefficients:
                      Estimate Std. Error z value Pr(>|z|)    
(Intercept)         -2.7671673  0.2500613 -11.066   <2e-16 ***
altitude_low_meters  0.0002810  0.0001872   1.501    0.133    
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

(Dispersion parameter for binomial family taken to be 1)

    Null deviance: 517.22  on 933  degrees of freedom
Residual deviance: 514.95  on 932  degrees of freedom
AIC: 518.95

Number of Fisher Scoring iterations: 5
Confusion Matrix and Statistics

          Reference
Prediction   0   1
         0 373  28
         1   0   0
                                          
               Accuracy : 0.9302          
                 95% CI : (0.9007, 0.9531)
    No Information Rate : 0.9302          
    P-Value [Acc > NIR] : 0.5501          
                                          
                  Kappa : 0               
                                          
 Mcnemar's Test P-Value : 3.352e-07       
                                          
            Sensitivity : 1.0000          
            Specificity : 0.0000          
         Pos Pred Value : 0.9302          
         Neg Pred Value :    NaN          
             Prevalence : 0.9302          
         Detection Rate : 0.9302          
   Detection Prevalence : 1.0000          
      Balanced Accuracy : 0.5000          
                                          
       'Positive' Class : 0               
                                          

From the previous, we can see that the logistic regression model is heavily biased towards predicting the majority class (class 0), with no instances of the minority class (class 1) being correctly predicted. This indicates that this model is no good to predict flavor score using altitude.

Random Forest

It seems that the flavor cannot be predicted using altitute in a linear manner. Thus, let’s try to use models that are not linear. In this case random forest.

Confusion Matrix and Statistics

          Reference
Prediction  0 6.08 6.17 6.33 6.42 6.5 6.58 6.67 6.75 6.83 6.92  7 7.08 7.17
      0     0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.08  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.17  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.33  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.42  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.5   0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.58  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.67  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.75  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      6.83  0    0    0    0    0   0    0    0    0    0    0  1    0    0
      6.92  0    0    0    0    0   0    0    0    0    0    0  0    0    1
      7     0    0    0    0    0   0    0    0    0    0    0  0    0    0
      7.08  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      7.17  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      7.25  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      7.33  0    0    0    0    0   1    0    0    0    0    1  0    0    2
      7.42  0    0    0    0    0   0    0    0    1    0    0  0    0    1
      7.5   0    0    0    0    0   0    1    0    0    1    2  2    2    4
      7.58  0    0    0    0    0   0    0    1    1    2    0  3    3    2
      7.67  0    0    0    0    0   0    0    0    0    0    0  1    0    0
      7.75  0    0    0    0    0   0    0    0    0    0    0  0    1    0
      7.81  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      7.83  0    0    0    0    0   0    0    0    0    0    0  0    0    1
      7.88  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      7.92  0    0    0    0    0   0    0    0    0    0    0  0    2    0
      8     0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.08  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.17  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.25  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.33  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.42  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.5   0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.58  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.67  0    0    0    0    0   0    0    0    0    0    0  0    0    0
      8.83  0    0    0    0    0   0    0    0    0    0    0  0    0    0
          Reference
Prediction 7.25 7.33 7.42 7.5 7.58 7.67 7.75 7.81 7.83 7.88 7.92  8 8.08 8.17
      0       0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.08    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.17    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.33    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.42    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.5     0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.58    0    0    0   0    0    1    0    0    0    0    0  0    0    0
      6.67    0    0    0   0    0    0    0    0    0    0    0  1    0    0
      6.75    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      6.83    0    0    0   0    0    0    0    0    0    0    1  0    0    0
      6.92    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      7       0    0    0   0    0    0    0    0    0    0    0  0    0    0
      7.08    0    1    0   1    1    0    1    0    0    0    0  0    0    0
      7.17    0    1    1   0    0    0    0    0    0    0    0  0    0    0
      7.25    0    0    0   0    1    0    0    0    0    0    0  0    0    0
      7.33    1    1    1   4    5    4    2    0    1    0    1  0    0    0
      7.42    0    1    3   4    2    2    1    0    2    0    2  1    0    0
      7.5     7    7    8  10    6    1    7    0    3    0    1  2    1    0
      7.58    3    7    5   9   12   10   12    0    9    0    4  3    0    3
      7.67    0    2    1   1    3    6    1    0    2    0    0  1    0    0
      7.75    1    0    2   3    3    5    0    0    0    0    0  0    0    0
      7.81    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      7.83    0    1    1   0    0    0    0    0    0    0    0  0    0    0
      7.88    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      7.92    0    0    0   1    0    0    1    0    0    0    0  0    1    0
      8       0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.08    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.17    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.25    0    1    0   0    0    0    0    0    0    0    0  0    0    0
      8.33    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.42    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.5     0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.58    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.67    0    0    0   0    0    0    0    0    0    0    0  0    0    0
      8.83    0    0    0   0    0    0    0    0    0    0    0  0    0    0
          Reference
Prediction 8.25 8.33 8.42 8.5 8.58 8.67 8.83
      0       0    0    0   0    0    0    0
      6.08    0    0    0   0    0    0    0
      6.17    0    0    0   0    0    0    0
      6.33    0    0    0   0    0    0    0
      6.42    0    0    0   0    0    0    0
      6.5     0    0    0   0    0    0    0
      6.58    0    0    0   0    0    0    0
      6.67    0    0    0   0    0    0    0
      6.75    0    0    0   0    0    0    0
      6.83    0    0    0   0    0    0    0
      6.92    0    0    0   0    0    0    0
      7       0    0    0   0    0    0    0
      7.08    0    0    0   0    0    0    0
      7.17    0    0    0   0    0    0    0
      7.25    0    0    0   0    0    0    0
      7.33    0    0    1   0    0    0    0
      7.42    0    0    0   0    0    0    0
      7.5     0    0    0   0    0    0    0
      7.58    1    0    0   1    0    0    0
      7.67    0    0    0   0    0    0    0
      7.75    0    0    0   0    0    0    0
      7.81    0    0    0   0    0    0    0
      7.83    0    0    0   0    0    0    0
      7.88    0    0    0   0    0    0    0
      7.92    0    0    0   0    0    0    0
      8       0    0    0   0    0    0    0
      8.08    0    0    0   0    0    0    0
      8.17    0    0    0   0    0    0    0
      8.25    0    0    0   0    0    0    0
      8.33    0    1    0   0    0    0    0
      8.42    0    0    0   0    0    0    0
      8.5     0    0    0   0    0    0    0
      8.58    0    0    0   0    0    0    0
      8.67    0    0    0   0    0    0    0
      8.83    0    0    0   0    0    0    0

Overall Statistics
                                          
               Accuracy : 0.1289          
                 95% CI : (0.0904, 0.1762)
    No Information Rate : 0.1289          
    P-Value [Acc > NIR] : 0.528           
                                          
                  Kappa : 0.0212          
                                          
 Mcnemar's Test P-Value : NA              

Statistics by Class:

                     Class: 0 Class: 6.08 Class: 6.17 Class: 6.33 Class: 6.42
Sensitivity                NA          NA          NA          NA          NA
Specificity                 1           1           1           1           1
Pos Pred Value             NA          NA          NA          NA          NA
Neg Pred Value             NA          NA          NA          NA          NA
Prevalence                  0           0           0           0           0
Detection Rate              0           0           0           0           0
Detection Prevalence        0           0           0           0           0
Balanced Accuracy          NA          NA          NA          NA          NA
                     Class: 6.5 Class: 6.58 Class: 6.67 Class: 6.75 Class: 6.83
Sensitivity            0.000000    0.000000    0.000000    0.000000    0.000000
Specificity            1.000000    0.996078    0.996078    1.000000    0.992095
Pos Pred Value              NaN    0.000000    0.000000         NaN    0.000000
Neg Pred Value         0.996094    0.996078    0.996078    0.992188    0.988189
Prevalence             0.003906    0.003906    0.003906    0.007812    0.011719
Detection Rate         0.000000    0.000000    0.000000    0.000000    0.000000
Detection Prevalence   0.000000    0.003906    0.003906    0.000000    0.007812
Balanced Accuracy      0.500000    0.498039    0.498039    0.500000    0.496047
                     Class: 6.92 Class: 7 Class: 7.08 Class: 7.17 Class: 7.25
Sensitivity             0.000000  0.00000     0.00000    0.000000    0.000000
Specificity             0.996047  1.00000     0.98387    0.991837    0.995902
Pos Pred Value          0.000000      NaN     0.00000    0.000000    0.000000
Neg Pred Value          0.988235  0.97266     0.96825    0.956693    0.952941
Prevalence              0.011719  0.02734     0.03125    0.042969    0.046875
Detection Rate          0.000000  0.00000     0.00000    0.000000    0.000000
Detection Prevalence    0.003906  0.00000     0.01562    0.007812    0.003906
Balanced Accuracy       0.498024  0.50000     0.49194    0.495918    0.497951
                     Class: 7.33 Class: 7.42 Class: 7.5 Class: 7.58 Class: 7.67
Sensitivity             0.045455     0.13636    0.30303     0.36364     0.20690
Specificity             0.897436     0.92735    0.75336     0.64574     0.94714
Pos Pred Value          0.040000     0.15000    0.15385     0.13187     0.33333
Neg Pred Value          0.909091     0.91949    0.87958     0.87273     0.90336
Prevalence              0.085938     0.08594    0.12891     0.12891     0.11328
Detection Rate          0.003906     0.01172    0.03906     0.04688     0.02344
Detection Prevalence    0.097656     0.07812    0.25391     0.35547     0.07031
Balanced Accuracy       0.471445     0.53186    0.52820     0.50469     0.57702
                     Class: 7.75 Class: 7.81 Class: 7.83 Class: 7.88
Sensitivity              0.00000          NA     0.00000          NA
Specificity              0.93506           1     0.98745           1
Pos Pred Value           0.00000          NA     0.00000          NA
Neg Pred Value           0.89627          NA     0.93281          NA
Prevalence               0.09766           0     0.06641           0
Detection Rate           0.00000           0     0.00000           0
Detection Prevalence     0.05859           0     0.01172           0
Balanced Accuracy        0.46753          NA     0.49372          NA
                     Class: 7.92 Class: 8 Class: 8.08 Class: 8.17 Class: 8.25
Sensitivity              0.00000  0.00000    0.000000     0.00000    0.000000
Specificity              0.97976  1.00000    1.000000     1.00000    0.996078
Pos Pred Value           0.00000      NaN         NaN         NaN    0.000000
Neg Pred Value           0.96414  0.96875    0.992188     0.98828    0.996078
Prevalence               0.03516  0.03125    0.007812     0.01172    0.003906
Detection Rate           0.00000  0.00000    0.000000     0.00000    0.000000
Detection Prevalence     0.01953  0.00000    0.000000     0.00000    0.003906
Balanced Accuracy        0.48988  0.50000    0.500000     0.50000    0.498039
                     Class: 8.33 Class: 8.42 Class: 8.5 Class: 8.58 Class: 8.67
Sensitivity             1.000000    0.000000   0.000000          NA          NA
Specificity             1.000000    1.000000   1.000000           1           1
Pos Pred Value          1.000000         NaN        NaN          NA          NA
Neg Pred Value          1.000000    0.996094   0.996094          NA          NA
Prevalence              0.003906    0.003906   0.003906           0           0
Detection Rate          0.003906    0.000000   0.000000           0           0
Detection Prevalence    0.003906    0.000000   0.000000           0           0
Balanced Accuracy       1.000000    0.500000   0.500000          NA          NA
                     Class: 8.83
Sensitivity                   NA
Specificity                    1
Pos Pred Value                NA
Neg Pred Value                NA
Prevalence                     0
Detection Rate                 0
Detection Prevalence           0
Balanced Accuracy             NA

Based on the previous, we managed to get rid of the bias from the previous model. However, the accuracy is too low. Hence, not making a very good model to predict flavor, at least just based on altitudes.

Random Forest with more features

Lets try to add more features that are not flavor related such as moisture and defects.

Confusion Matrix and Statistics

          Reference
Prediction 0 6.08 6.17 6.33 6.42 6.5 6.58 6.67 6.75 6.83 6.92 7 7.08 7.17 7.25
      0    0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.08 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.17 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.33 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.42 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.5  0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.58 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.67 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.75 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      6.83 0    0    0    0    0   0    0    0    0    0    0 0    0    0    1
      6.92 0    0    0    0    0   0    0    0    0    0    0 0    0    1    0
      7    0    0    0    0    0   0    0    0    1    0    0 1    2    0    1
      7.08 0    0    0    0    0   0    0    0    0    0    0 0    1    0    0
      7.17 0    0    0    0    0   0    0    0    0    1    1 0    0    1    0
      7.25 0    0    0    0    0   0    0    0    0    0    0 0    0    1    0
      7.33 0    0    0    0    0   0    1    0    0    1    1 2    1    0    2
      7.42 0    0    0    0    0   0    0    0    0    0    0 0    1    0    0
      7.5  0    0    0    0    0   0    0    0    0    0    1 1    0    2    4
      7.58 0    0    0    0    0   0    0    0    0    1    0 1    3    0    0
      7.67 0    0    0    0    0   0    0    0    0    0    0 0    0    2    2
      7.75 0    0    0    0    0   0    0    1    0    0    0 1    0    1    1
      7.81 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      7.83 0    0    0    0    0   0    0    0    1    0    0 0    0    1    0
      7.88 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      7.92 0    0    0    0    0   0    0    0    0    0    0 0    0    1    0
      8    0    0    0    0    0   1    0    0    0    0    0 1    0    0    1
      8.08 0    0    0    0    0   0    0    0    0    0    0 0    0    1    0
      8.17 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.25 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.33 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.42 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.5  0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.58 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.67 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
      8.83 0    0    0    0    0   0    0    0    0    0    0 0    0    0    0
          Reference
Prediction 7.33 7.42 7.5 7.58 7.67 7.75 7.81 7.83 7.88 7.92 8 8.08 8.17 8.25
      0       0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.08    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.17    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.33    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.42    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.5     0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.58    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.67    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.75    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.83    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      6.92    0    0   0    0    0    1    0    0    0    0 0    0    0    0
      7       1    0   0    0    0    0    0    0    0    0 0    0    0    0
      7.08    1    1   0    1    1    0    0    1    0    0 0    0    0    0
      7.17    1    1   0    2    0    2    0    0    0    1 0    0    0    0
      7.25    0    0   2    3    0    3    0    1    0    1 0    0    0    0
      7.33    0    1   3    3    5    0    0    1    0    0 0    0    0    0
      7.42    2    1   1    1    1    1    0    1    0    2 2    0    0    1
      7.5     7    4   4    7    4    6    0    4    0    2 1    0    1    0
      7.58    7    5   8    6    8    5    0    4    0    0 2    1    1    0
      7.67    3    4   6    5    2    4    0    0    0    1 1    0    0    0
      7.75    0    4   8    4    3    3    0    3    0    1 1    0    0    0
      7.81    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      7.83    0    1   0    1    2    0    0    2    0    1 1    1    0    0
      7.88    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      7.92    0    0   1    0    1    0    0    0    0    0 0    0    0    0
      8       0    0   0    0    2    0    0    0    0    0 0    0    1    0
      8.08    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.17    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.25    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.33    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.42    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.5     0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.58    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.67    0    0   0    0    0    0    0    0    0    0 0    0    0    0
      8.83    0    0   0    0    0    0    0    0    0    0 0    0    0    0
          Reference
Prediction 8.33 8.42 8.5 8.58 8.67 8.83
      0       0    0   0    0    0    0
      6.08    0    0   0    0    0    0
      6.17    0    0   0    0    0    0
      6.33    0    0   0    0    0    0
      6.42    0    0   0    0    0    0
      6.5     0    0   0    0    0    0
      6.58    0    0   0    0    0    0
      6.67    0    0   0    0    0    0
      6.75    0    0   0    0    0    0
      6.83    0    0   0    0    0    0
      6.92    0    0   0    0    0    0
      7       0    0   0    0    0    0
      7.08    0    0   0    0    0    0
      7.17    0    0   0    0    0    0
      7.25    0    0   0    0    0    0
      7.33    0    0   0    0    0    0
      7.42    0    0   0    0    0    0
      7.5     0    0   0    0    0    0
      7.58    0    0   0    0    0    0
      7.67    1    0   0    0    0    0
      7.75    0    1   0    0    0    0
      7.81    0    0   0    0    0    0
      7.83    0    0   0    0    0    0
      7.88    0    0   0    0    0    0
      7.92    0    0   0    0    0    0
      8       0    0   0    0    0    0
      8.08    0    0   0    0    0    0
      8.17    0    0   0    0    0    0
      8.25    0    0   0    0    0    0
      8.33    0    0   0    0    0    0
      8.42    0    0   0    0    0    0
      8.5     0    0   0    0    0    0
      8.58    0    0   0    0    0    0
      8.67    0    0   1    0    0    0
      8.83    0    0   0    0    0    0

Overall Statistics
                                          
               Accuracy : 0.082           
                 95% CI : (0.0515, 0.1227)
    No Information Rate : 0.1289          
    P-Value [Acc > NIR] : 0.993           
                                          
                  Kappa : -0.0169         
                                          
 Mcnemar's Test P-Value : NA              

Statistics by Class:

                     Class: 0 Class: 6.08 Class: 6.17 Class: 6.33 Class: 6.42
Sensitivity                NA          NA          NA          NA          NA
Specificity                 1           1           1           1           1
Pos Pred Value             NA          NA          NA          NA          NA
Neg Pred Value             NA          NA          NA          NA          NA
Prevalence                  0           0           0           0           0
Detection Rate              0           0           0           0           0
Detection Prevalence        0           0           0           0           0
Balanced Accuracy          NA          NA          NA          NA          NA
                     Class: 6.5 Class: 6.58 Class: 6.67 Class: 6.75 Class: 6.83
Sensitivity            0.000000    0.000000    0.000000    0.000000    0.000000
Specificity            1.000000    1.000000    1.000000    1.000000    0.996047
Pos Pred Value              NaN         NaN         NaN         NaN    0.000000
Neg Pred Value         0.996094    0.996094    0.996094    0.992188    0.988235
Prevalence             0.003906    0.003906    0.003906    0.007812    0.011719
Detection Rate         0.000000    0.000000    0.000000    0.000000    0.000000
Detection Prevalence   0.000000    0.000000    0.000000    0.000000    0.003906
Balanced Accuracy      0.500000    0.500000    0.500000    0.500000    0.498024
                     Class: 6.92 Class: 7 Class: 7.08 Class: 7.17 Class: 7.25
Sensitivity             0.000000 0.142857    0.125000    0.090909     0.00000
Specificity             0.992095 0.979920    0.979839    0.963265     0.95492
Pos Pred Value          0.000000 0.166667    0.166667    0.100000     0.00000
Neg Pred Value          0.988189 0.976000    0.972000    0.959350     0.95102
Prevalence              0.011719 0.027344    0.031250    0.042969     0.04688
Detection Rate          0.000000 0.003906    0.003906    0.003906     0.00000
Detection Prevalence    0.007812 0.023438    0.023438    0.039062     0.04297
Balanced Accuracy       0.496047 0.561388    0.552419    0.527087     0.47746
                     Class: 7.33 Class: 7.42 Class: 7.5 Class: 7.58 Class: 7.67
Sensitivity              0.00000    0.045455    0.12121     0.18182    0.068966
Specificity              0.91026    0.944444    0.80269     0.79372    0.872247
Pos Pred Value           0.00000    0.071429    0.08333     0.11538    0.064516
Neg Pred Value           0.90638    0.913223    0.86058     0.86765    0.880000
Prevalence               0.08594    0.085938    0.12891     0.12891    0.113281
Detection Rate           0.00000    0.003906    0.01562     0.02344    0.007812
Detection Prevalence     0.08203    0.054688    0.18750     0.20312    0.121094
Balanced Accuracy        0.45513    0.494949    0.46195     0.48777    0.470606
                     Class: 7.75 Class: 7.81 Class: 7.83 Class: 7.88
Sensitivity              0.12000          NA    0.117647          NA
Specificity              0.87446           1    0.962343           1
Pos Pred Value           0.09375          NA    0.181818          NA
Neg Pred Value           0.90179          NA    0.938776          NA
Prevalence               0.09766           0    0.066406           0
Detection Rate           0.01172           0    0.007812           0
Detection Prevalence     0.12500           0    0.042969           0
Balanced Accuracy        0.49723          NA    0.539995          NA
                     Class: 7.92 Class: 8 Class: 8.08 Class: 8.17 Class: 8.25
Sensitivity              0.00000  0.00000    0.000000     0.00000    0.000000
Specificity              0.98785  0.97581    0.996063     1.00000    1.000000
Pos Pred Value           0.00000  0.00000    0.000000         NaN         NaN
Neg Pred Value           0.96443  0.96800    0.992157     0.98828    0.996094
Prevalence               0.03516  0.03125    0.007812     0.01172    0.003906
Detection Rate           0.00000  0.00000    0.000000     0.00000    0.000000
Detection Prevalence     0.01172  0.02344    0.003906     0.00000    0.000000
Balanced Accuracy        0.49393  0.48790    0.498031     0.50000    0.500000
                     Class: 8.33 Class: 8.42 Class: 8.5 Class: 8.58 Class: 8.67
Sensitivity             0.000000    0.000000   0.000000          NA          NA
Specificity             1.000000    1.000000   1.000000           1    0.996094
Pos Pred Value               NaN         NaN        NaN          NA          NA
Neg Pred Value          0.996094    0.996094   0.996094          NA          NA
Prevalence              0.003906    0.003906   0.003906           0    0.000000
Detection Rate          0.000000    0.000000   0.000000           0    0.000000
Detection Prevalence    0.000000    0.000000   0.000000           0    0.003906
Balanced Accuracy       0.500000    0.500000   0.500000          NA          NA
                     Class: 8.83
Sensitivity                   NA
Specificity                    1
Pos Pred Value                NA
Neg Pred Value                NA
Prevalence                     0
Detection Rate                 0
Detection Prevalence           0
Balanced Accuracy             NA

By adding more features, we actually did worse. Because of this, we can assume that there is no enough data to make a good use of random forest.

Result, Analysis, and Discussion

The logistic regression models demonstrated varying degrees of effectiveness in predicting coffee flavor scores based on different sets of features:

Results

  1. Single Feature (Aroma):

    Accuracy: 93.53%

    Sensitivity: 98.13%

    Specificity: 32.14%

    Positive Predictive Value: 95.08%

    Kappa: 0.3776

  2. Multiple Features (Aroma, Aftertaste, Acidity, Body, Balance, Uniformity, Clean Cup, Sweetness):

    Accuracy: 96.27%

    Sensitivity: 98.93%

    Specifity: 60.71%

    Positive Predictive Value: 97.11%

    Kappa: 0.6744

    (Improved metrics across the board, indicating a more robust model)

  3. External Features (Altitude):

    The model failed to predict the minority class (Flavor score of 1), highlighting the inadequacy of altitude as a predictive feature for coffee flavor.

  4. Random Forest

    The model accuracy is incredibly low. This indicates that the model is under performing. Thus, in order to improve the model more data is needed.

Analysis

The initial exploratory data analysis (EDA) revealed several important insights:

  1. Species Distribution: The majority of coffee samples are Arabica, suggesting a focus on testing this species.

  2. Country of Origin Distribution: Mexico, Guatemala, and Colombia are the top contributors to the dataset, indicating a concentration of coffee quality testing in these regions.

  3. Altitude Distribution: Most coffee samples are grown at an average altitude of 1200 meters, aligning with known optimal conditions for high-quality coffee.

  4. Total Cup Points Distribution: The overall high scores across countries suggest a generally high quality of tested coffee samples, with notable outliers like Honduras with a score of 0, indicating potential data anomalies.

  5. Correlation Matrix: Strong correlations were observed among flavor profile attributes, indicating the possibility of predicting one flavor score based on others. Defects and altitude showed little to no correlation with flavor scores, suggesting limited predictive power for these features.

  6. Linear Models: Linear model seems to perform very well for features with strong correlation in this case predicting flavor score based on other flavor profiles.

  7. Non-linear Models: Random forest did not perform well when trying to predict the flavor score based on features with low correlation in this case non-flavor based features (altitudes, moisture, defects). These could indicate that we lack of data in order to accurately predict flavor score based on non-flavor information.

Discussion

Predictive Modeling:

The logistic regression model using Aroma as the sole predictor of Flavor achieved high accuracy and sensitivity but struggled with specificity. This indicates that while Aroma is a strong indicator of coffee flavor, it may not be sufficient on its own for comprehensive prediction.

Adding multiple flavor-related features significantly improved the model’s performance, highlighting the importance of a multifaceted approach when predicting complex attributes like coffee flavor.

Using external features such as altitude for flavor prediction proved ineffective. This aligns with the observed lack of correlation between altitude and flavor scores, suggesting that altitude alone does not capture the nuances influencing coffee flavor.

Implications:

The strong inter-correlation among flavor attributes implies that coffee quality assessments are inherently interconnected. Models predicting coffee ratings or specific flavor scores should incorporate multiple related features to enhance accuracy.

The data suggests a geographical bias towards certain coffee-producing countries. This could reflect the industry’s focus on these regions or data collection biases.

The presence of outliers and anomalies, such as the zero scores, indicates a need for further data cleaning and validation to ensure the reliability of predictive models.

Future Work:

Exploring other non-linear models and more complex machine learning techniques may yield better predictions, especially when incorporating diverse features like processing methods and geographical data. Additional data collection from underrepresented regions and varieties could provide a more balanced dataset, enhancing model generalizability. Investigating other external factors, such as soil quality, weather conditions, and farming practices, could offer deeper insights into the determinants of coffee quality.

Impact

This study could have several significant impacts on the coffee industry and Sustainability and Ethical Considerations:

  1. Coffee Industry:
  • Improved Quality Control: The findings from this study can help coffee producers and quality control teams identify the key attributes that contribute to high coffee ratings. By understanding the significance of factors such as aroma, aftertaste, and balance, producers can refine their processes to enhance these characteristics, thereby improving the overall quality of their coffee.
  • Enhanced Marketability: Coffee brands can use these insights to market their products more effectively. Highlighting specific flavor attributes and the conditions under which their coffee is grown can create compelling narratives that attract consumers who are looking for high-quality coffee. This can differentiate products in a competitive market and build brand loyalty.
  • Geographical Insights: The analysis identifies regions that consistently produce high-quality coffee. This information can guide investment decisions and resource allocation for coffee producers and traders. Promoting regions with a track record of excellence can enhance their reputation and market share, while discovering new high-quality coffee sources can expand the diversity of offerings.
  1. Sustainability and Ethical Considerations:
  • Sustainable Practices: By pinpointing the key factors that lead to high-quality coffee, the study can guide sustainable agricultural practices. Coffee producers can optimize these factors to produce high-quality coffee with minimal environmental impact, promoting sustainability within the industry. This approach supports the long-term viability of coffee farming and helps preserve the environment.

  • Data Transparency and Accessibility: The use of publicly available data sets a precedent for transparency and accessibility in research. By making data accessible, this study encourages other researchers to use and share data, fostering collaboration and innovation in the field. This openness can lead to more comprehensive studies and a deeper understanding of coffee quality, benefiting the entire coffee industry and consumers alike.