Introduction

I found my dataset when I came across the the ‘Coasters_2015’ set on the DASL site. This dataset contains information on more than 200 roller coasters that opened in 2015. The original dataset includes variables such as Name, Park, Track material (steel or wood), Speed, Height, Length, Drop, Duration, and Inversions. After reviewing the dataset, however, I noticed that many entries were missing data for the Drop and Duration categories. I decided to remove these two variables from the dataset before moving forward with my analysis. I also added a new variable to my dataset—Country, which I obtained from the original source of data, the Roller Coaster Database (RCDB).

The RCDB collects its data using a largely community driven-approach, in which roller coaster enthusiasts and users of the site submit information and photos of roller coasters they visit. This information is verified through publicly available sources such as park websites, press releases, and manufacturer reports, and direct communications with parks and coaster manufacturers before being compiled into the database. While the data does undergo cross-validation, it is worth noting that it does rely partly on user submissions which can lead to gaps in information as I have seen in the Drop and Duration categories of my dataset. It is also important to note that my data does not come from a random sample. It is exclusively made up of roller coasters that were opened in 2015, and because of this the data cannot be used to make general statements about roller coasters across time. However, it provides insight to trends in roller coaster construction during 2015.

My final dataset includes 240 observations of the following variables: Park, Country, Track, Speed, Height, Length, and Inversions. I removed the Name variable as it does not provide meaningful information for analysis. The variable definitions are as follows:

Park describes the amusement park where the roller coaster is located. Country describes the country where the roller coaster was built. Track indicates whether the roller coaster’s track is made of steel or wood. Speed, Height, and Length are quantitative variables describing the coaster’s top speed (in mph), maximum height (in feet), and track length (in feet), respectively. Inversions describes the number of inversions a roller coaster features. However, since no roller coaster in this dataset has more than one inversion, I treated it as a binary categorical variable (whether or not a coaster has an inversion).

The main goal of my analysis is to investigate the relationship between the speed of roller coasters and other factors such as the height, length, type of track, and whether or not it features an inversion. I am curious to discover which factors are the most predictive of roller coaster speed. I would then like to use my model to predict the speed of roller coasters with certain specifications. In order to perform my analysis I am using Lasso regression, a form of regularized regression and shrinkage method.

Methodology

Lasso, Least Absolute Shrinkage and Selection Operator, regression is an extension of Ordinary Least Squares (OLS) regression that utilizes something called L1 regularization. A penalty term is added to the function that minimizes the sum of squared residuals, which helps to prevent overfitting and avoid overly complex models. Lasso can be a method of model selection, and behaves in a hypothesis-generating manner rather than hypothesis-testing like standard linear regression.

Lasso regression minimizes this objective function:

\[\frac{1}{2N}\sum_{i=1}^{N} (y_i - \beta_0 - \sum_{j=1}^{p} x_{ij}\beta_j)^2 + \lambda \sum_{j=1}^{p} |\beta_j|\]

or, simply put, \(RSS + L_1\).

The first term in this function, \(\frac{1}{2N}\sum_{i=1}^{N}(y_i-\beta_0-\sum_{j=1}^{p}x_{ij}\beta_j)^2\), is recognizable from ordinary least squares regression. \(\sum_{i=1}^{N}(y_i-\beta_0-\sum_{j=1}^{p}x_{ij}\beta_j)^2\) is the sum of squared residuals or the error term. \(\frac{1}{2N}\) is the scaling factor applied to the error term, which averages the errors across the number of observations to normalize them and simplify the calculations in the optimization process.

The second term in the objective function is the penalty term, also called the L1 Regularization term. \(L_1=\lambda \sum_{j=1}^{p} |\beta_j|\) is the sum of the absolute values of the coefficients multiplied by \(\lambda\), a tuning parameter. The value of \(\lambda\) controls the amount of regularization, or penalty, that is applied. The penalty term helps to prevent overfitting, and is what allows Lasso regression to act as a means of model selection. Rather than selecting predictors for a model by hand based on speculation about what variables will be the most predictive, Lasso regression assesses the contribution of all predictors to the model, and the application of the penalty term shrinks less-predictive coefficients accordingly. If a predictor is not statistically significant and does not contribute to the model, Lasso is able to shrink its coefficient entirely to zero, effectively eliminating it from the model. The tuning parameter, \(\lambda\), can range in value from zero to one. When \(\lambda=0\), the model does not apply regularization and performs a standard linear regression. As values of \(\lambda\) increase, the model becomes more regularized and more coefficients are shrunk towards or to zero. \(\lambda=1\) would lead to a model with no predictors. The optimal value of \(\lambda\) is determined through k-fold cross validation.

Another form of regularized regression or shrinkage method is Ridge regression, which features L2 regularization. The penalty term in the objective function of Ridge regression differs slightly from that of Lasso regression, it is \(L_2=\lambda \sum_{j=1}^{p} \beta^2_j\), the sum of the squared coefficients. Because of this, Ridge regression is able to shrink coefficients towards zero, but is not able to eliminate predictors from the model. Because of this, Lasso is preferable in situations with the goal of identifying a subset of important predictors from a larger group.

Lasso and Ridge regression can also be understood as constrained optimization problems using Lagrangian multipliers. They begin with a standard linear regression model in which \(y_i=\beta_0+\beta_1X_1+...+\beta_nX_n + \epsilon\). In an unconstrained optimization or standard linear regression, one would simply minimize \(\sum_{i=1}^{N} (y_i - \beta_0 - \sum_{j=1}^{p} x_{ij}\beta_j)^2\). To apply the L1 regularization, however, the function must be minimized subject to the penalty term scaled by the Lagrangian multiplier, or tuning parameter, \(\lambda\).

Lasso regression has certain conditions that must be met, similar to those of standard linear regression. Key assumptions include linearity of relationships between predictors and the response variable, independence of observations, no perfect multicollinearity, and standardization of variables.

OLS has an assumption of homoskedasticity or constant variance of errors. this assumption does not translate into Lasso regression. Heteroskedasticity does not lead to bias in Lasso estimates, but may effect the precision of the estimates. If heteroskedasticity is suspected, one may consider using a weighted approach to reduce impact on the final model.

Lasso regression can handle certain degrees of multicollinearity more easily than standard linear regression, but in instances of perfect multicollinearity Lasso tends to behave unreliably by arbitrarily selecting one predictor and shrinking the others to zero.

Standardization of variables is useful when performing Lasso or other regularized regression such as Ridge or Elastic Net. This is because regularized regression can be sensitive to the scale of variable inputs. Without standardization, predictors with larger scales may have disproportionate influence on the final model. The standardization process, which is applied automatically when using the glmnet package in R, prevents large-scale variables from dominating the model. The intercept term is unaffected by this process.

In order to avoid overfitting, it is recommended that a dataset be randomly split into “Train” and “Test” groups. The training group contains the bulk of the observations, typically 70-80%, and is used to build the model. The remaining observations become the test group, which is set aside and is used to evaluate the model once it has been trained, allowing for assessment of the model’s ability to generalize to data it has not seen previously. The test group also validates the choice of \(\lambda\) on independent data, preventing inadvertent bias from evaluating the model on the same data with which it was trained.

library(glmnet)
## Loaded glmnet 4.1-8
coasters<- read_csv("spfinaldata.csv", show_col_types = FALSE)

I begin by converting the categorical variable Track into a binary dummy variable, in which the value one indicates a steel track and zero indicates wooden. I also note that the variable Inversions only shows values zero and one, as none of the roller coasters built in 2015 contained in this dataset feature more than one inversion. As such, I will treat the variable Inversions as a binary dummy variable with the understanding that results will likely not be applicable to coasters that feature more than one inversion.

coasters<- coasters %>% 
  mutate(Track=ifelse(Track=="Steel",1,0))

Plotting the relationships between my response variable, Speed, and my two numeric explanatory variables, Length and Height, allows me to check for linearity. I also generate boxplots to visualize the relationship between my response variable and my binary categorical variables, Track and Inversions, and numeric summaries of the relationships between Speed and the Park and Country variables, which are categorical but not binary.

xyplot(Speed~Length, col="pink", data=coasters)

xyplot(Speed~Height, col="seagreen1", data=coasters)

boxplot(Speed~Inversions, col="plum", data=coasters)

boxplot(Speed~Track, col="skyblue1", data=coasters)

favstats(~Speed, group=Country, data=coasters)[c("Country","mean","sd", "n")]
##        Country      mean         sd   n
## 1    Australia  52.80000         NA   1
## 2      Austria  23.95000  2.1920310   2
## 3      Belgium  45.05000 15.3442172   2
## 4     Botswana  47.90000         NA   1
## 5       Brazil  17.40000         NA   1
## 6       Canada  46.92500 14.2626260   4
## 7        China  53.00811 11.2926815  37
## 8   Costa Rica  47.00000         NA   1
## 9      Denmark  40.40000         NA   1
## 10     Finland  65.10000  0.2828427   2
## 11      France  37.31429 11.5668245   7
## 12     Germany  55.71000 14.4991532  10
## 13       India  60.86000 26.6473826   5
## 14   Indonesia  62.10000         NA   1
## 15        Iraq  29.10000         NA   1
## 16     Ireland  55.90000         NA   1
## 17       Italy  51.67778 13.6995235   9
## 18       Japan  62.88000 20.9448907  10
## 19      Mexico  62.00000         NA   1
## 20     Morocco  43.50000         NA   1
## 21 Netherlands  43.50000  0.0000000   2
## 22 North Korea  25.70000         NA   1
## 23      Poland  41.60000 11.4503275   3
## 24     Romania  21.80000         NA   1
## 25      Russia  62.10000         NA   1
## 26       Spain  58.96667 28.2404556   3
## 27      Sweden  30.24000 21.3156750   5
## 28 Switzerland  52.80000         NA   1
## 29      Taiwan  67.10000         NA   1
## 30      Turkey  55.95000 17.6069589   2
## 31         UAE 149.10000         NA   1
## 32          UK  34.84000 18.6116361   5
## 33         USA  60.51565 16.6872343 115
## 34     Vietnam  21.70000         NA   1
favstats(~Speed, group=Park, data=coasters)[c("Park","mean","sd", "n")]
##                                        Park      mean        sd  n
## 1                            Adlabs Imagica  54.95000 14.495689  2
## 2                            Adventure City  31.10000        NA  1
## 3                           Adventure World  52.80000        NA  1
## 4                             Adventuredome  43.00000  2.828427  2
## 5                              Alton Towers  68.00000        NA  1
## 6                                 Bagatelle  39.55000 14.778532  2
## 7                               Bayern Park  49.70000        NA  1
## 8                                  Belantis  52.80000        NA  1
## 9                                 Big Sheep  26.80000        NA  1
## 10                            Bobbejaanland  34.20000        NA  1
## 11                            Bosque Magico  62.00000        NA  1
## 12           Buffalo Bill's Resort & Casino  80.00000        NA  1
## 13                      Busch Gardens Tampa  60.00000        NA  1
## 14               Busch Gardens Williamsburg  67.33333  5.507571  3
## 15                        Canobie Lake Park  43.50000        NA  1
## 16                                Carowinds  72.33333 24.110855  3
## 17                          Cavallino Matto  51.00000        NA  1
## 18                              Cedar Point  76.16667 27.345323  6
## 19                   Century Amusement Park  49.70000        NA  1
## 20                  Chimelong Ocean Kingdom  67.10000        NA  1
## 21                  Chuanlord Holiday Manor  65.30000        NA  1
## 22                   Cliff's Amusement Park  47.00000        NA  1
## 23                               Conny-Land  52.80000        NA  1
## 24                       Crab Island Resort  49.70000        NA  1
## 25               Cultus Lake Adventure Park  26.00000        NA  1
## 26                                Dollywood  63.00000        NA  1
## 27                              Dorney Park  70.00000  7.071068  2
## 28          Dorney Park & Wildwater Kingdom  50.00000        NA  1
## 29                               Dream City  43.50000        NA  1
## 30                                Dreamland  29.10000        NA  1
## 31                               Dreamworld  82.65000 24.536605  2
## 32                                 Duinrell  43.50000        NA  1
## 33                          E-DA Theme Park  67.10000        NA  1
## 34                             Energylandia  41.60000 11.450328  3
## 35                  Erlebnispark Tripsdrill  62.10000        NA  1
## 36                                Euro Park  46.60000        NA  1
## 37                              Europa Park  62.10000  0.000000  2
## 38                                 Europark  29.10000        NA  1
## 39                      Fantawild Adventure  56.15000 12.940054  2
## 40                     Fantawild Dream Park  51.93333 11.707405  3
## 41                      Fantawild Dreamland  65.30000        NA  1
## 42                  Ferrari World Abu Dhabi 149.10000        NA  1
## 43           Flamingo Land Theme Park & Zoo  25.70000        NA  1
## 44                                Floraland  46.60000        NA  1
## 45                            Frontier City  55.00000        NA  1
## 46                          Fuji-Q Highland  62.10000        NA  1
## 47                         Fuji-Q Highlands  80.80000        NA  1
## 48                                 Fun Spot  45.00000        NA  1
## 49                             Fun Spot USA  29.10000        NA  1
## 50                                  Funplex  26.70000        NA  1
## 51                Funtown Splashtown U.S.A.  28.00000        NA  1
## 52  Galveston Island Historic Pleasure Pier  52.00000        NA  1
## 53                                Gardaland  59.00000  4.384062  2
## 54               Giant Wheel Park of Suzhou  46.60000        NA  1
## 55          Glenwood Caverns Adventure Park  16.20000        NA  1
## 56                               Grona Lund  20.90000 23.193102  2
## 57                    Gulliver's Warrington  24.60000        NA  1
## 58                        Hangzhou Paradise  49.70000        NA  1
## 59                               Hansa Park  78.90000        NA  1
## 60                             Happy Valley  57.16000 15.035724  5
## 61                              Happy World  44.70000        NA  1
## 62                    Harbin Amusement Park  49.70000        NA  1
## 63                        Heide Park Resort  62.10000        NA  1
## 64                              Hersheypark  62.50000 17.677670  2
## 65                             Holiday Park  45.00000 24.041631  2
## 66                            Holiday World  60.00000        NA  1
## 67                                  HotZone  17.40000        NA  1
## 68                    Jin Jiang Action Park  65.60000        NA  1
## 69                      Jinling Happy World  49.70000        NA  1
## 70                        Kaeson Youth Park  25.70000        NA  1
## 71                                Kennywood  50.00000        NA  1
## 72                           Kennywood Park  72.33333 15.044379  3
## 73                           Kings Dominion  90.00000        NA  1
## 74                             Kings Island  74.00000  8.485281  2
## 75         Knoebels Amusement Park & Resort  55.00000        NA  1
## 76                       Knott's Berry Farm  82.00000        NA  1
## 77                                Kolmarden  24.90000        NA  1
## 78                                 La Ronde  49.70000        NA  1
## 79                                   Lagoon  70.00000        NA  1
## 80                       Lake Winnepesaukah  50.00000        NA  1
## 81                                    LaQua  80.80000        NA  1
## 82                         Legoland Billund  40.40000        NA  1
## 83                         Legoland Florida  35.00000        NA  1
## 84                           Lewa Adventure  71.50000        NA  1
## 85                               Linnanmaki  65.30000        NA  1
## 86                         Lion Park Resort  47.90000        NA  1
## 87                                 Liseberg  62.10000        NA  1
## 88              Loca Joy Holiday Theme Park  52.80000        NA  1
## 89                                 Lunapark  40.00000        NA  1
## 90                             Mer de Sable  28.00000        NA  1
## 91                             Mirabilandia  74.60000        NA  1
## 92             Miracle Strip Amusement Park  55.00000        NA  1
## 93                                 Miragica  28.60000        NA  1
## 94                           Movieland Park  50.00000        NA  1
## 95                    Myrtle Beach Pavilion  55.00000        NA  1
## 96                       Nagashima Spa Land  75.45000 27.647875  2
## 97                   Nantong Adventure Land  65.30000        NA  1
## 98        New York, New York Hotel & Casino  67.00000        NA  1
## 99                        Oriental Heritage  49.76667  4.792007  3
## 100           Paramount Canada's Wonderland  56.00000  0.000000  2
## 101               Paramount's Great America  55.00000        NA  1
## 102              Paramount's Kings Dominion  75.00000  7.071068  2
## 103                Paramount's Kings Island  71.60000  9.616652  2
## 104                         Parc du Bocasse  29.10000        NA  1
## 105         Parque de Atracciones de Madrid  28.00000        NA  1
## 106                      Parque Diversiones  47.00000        NA  1
## 107                           People's Park  20.50000        NA  1
## 108                         Planet Infiniti  29.10000        NA  1
## 109                     Plopsaland De Panne  55.90000        NA  1
## 110                       PortAventura Park  83.30000        NA  1
## 111                               PowerLand  64.90000        NA  1
## 112                           Puyallup Fair  50.00000        NA  1
## 113                           Race City PCB  35.00000        NA  1
## 114                       Rainbow MagicLand  47.50000 16.263456  2
## 115                             Scream Zone  25.70000        NA  1
## 116                       SeaWorld Of Texas  65.00000        NA  1
## 117                        SeaWorld Orlando  60.50000  6.363961  2
## 118               Shenyang Botanical Garden  49.70000        NA  1
## 119                      Silver Dollar City  67.00000  1.414214  2
## 120                   Silverwood Theme Park  46.00000        NA  1
## 121                             Sinbad Land  29.10000        NA  1
## 122                                Sindibad  43.50000        NA  1
## 123                         Sino Wonderland  49.70000        NA  1
## 124                       Six Flags America  54.33333  1.154701  3
## 125                   Six Flags Darien Lake  73.00000        NA  1
## 126             Six Flags Discovery Kingdom  62.00000        NA  1
## 127                  Six Flags Fiesta Texas  61.50000 16.010413  4
## 128               Six Flags Great Adventure  70.33333  8.736895  3
## 129                 Six Flags Great America  60.32000 12.091815  5
## 130              Six Flags Kentucky Kingdom  59.00000  5.656854  2
## 131                Six Flags Magic Mountain  66.51000 17.736369 10
## 132                  Six Flags Marine World  60.00000  7.071068  2
## 133                   Six Flags New England  56.40000 20.956463  4
## 134                  Six Flags Over Georgia  52.00000        NA  1
## 135                    Six Flags Over Texas  75.00000 14.142136  2
## 136                     Six Flags St. Louis  58.43333 10.132292  3
## 137           Six Flags Worlds of Adventure  65.00000        NA  1
## 138                        Skara Sommarland  22.40000        NA  1
## 139                            Skyline Park  37.30000        NA  1
## 140                             Space World  71.50000        NA  1
## 141                   Suzhou Amusement Land  49.70000        NA  1
## 142                              Tayto Park  55.90000        NA  1
## 143                  Tianjin Shuishang Park  49.70000        NA  1
## 144                          Tokyo Joypolis  23.60000        NA  1
## 145                        tokyo SummerLand  60.30000        NA  1
## 146                        Tokyo SummerLand  60.30000        NA  1
## 147                               Toverland  43.50000        NA  1
## 148          Trans Studio Indoor Theme Park  62.10000        NA  1
## 149  Universal Studios Islands of Adventure  67.00000        NA  1
## 150                             Valleyfair!  74.00000        NA  1
## 151                                  ViaSea  55.95000 17.606959  2
## 152                           Vinpearl Land  21.70000        NA  1
## 153                          Walygator Parc  55.90000        NA  1
## 154                Warner Bros. Movie World  65.60000        NA  1
## 155                   Washington State Fair  45.00000        NA  1
## 156                              Weihe Park  49.70000        NA  1
## 157                           Wiener Prater  23.95000  2.192031  2
## 158                         Wild Adventures  53.50000  2.121320  2
## 159                           Wonder Island  62.10000        NA  1
## 160                           World Joyland  60.00000  7.495332  2
## 161                           Worlds of Fun  64.33333 11.015141  3
## 162                             Yexian Park  32.90000        NA  1
## 163                             Yomiuriland  38.50000        NA  1
## 164                    ZDT'S Amusement Park  40.00000        NA  1
## 165                               Zoomarine  47.90000        NA  1
## 166                          Zorky's Planet  21.80000        NA  1

We can see that there is a positive linear relationship between Speed and Length, as well as Speed and Height, in which Speed increases as the Length or Height increases. The first boxplot shows that roller coasters with zero and one inversions have a similar mean speed, but roller coasters with an inversion have a smaller range of speeds. Both groups have a small quantity of outliers, the group with no inversions has an outlier at high speeds, and the group with one inversion has two outliers at low speeds. The second boxplot shows that roller coasters with a wooden track (indicated by the value zero) have a similar mean speed to roller coasters with a steel track (indicated by the value one), but those with wooden tracks face a smaller range of speeds. The steel groups has a few outliers, most at high speeds. The favstats output provides more information than is easily digestible, some notable pieces are that: The U.S. and China built the most roller coasters in 2015, with 115 and 37 respectively. The fastest roller coaster was built in the UAE and reaches speeds of 149.1 mph. The amusement parks that introduced the most new roller coasters in 2015 are Six Flags Magic Mountain and Cedar Point, with 10 and 6 coasters respectively. I will not be utilizing the Park or Country variables in my regression, as they are neither quantitative nor binary categorical variables and would be too difficult to work with.

Limiting my dataset to quantitative and binary categorical variables, I am able to check for multicollinearity amongst my data.

coasters.num<- coasters[,-1:-2]
cor(coasters.num)
##                  Track       Speed      Height     Length   Inversions
## Track       1.00000000 -0.05175766 0.081983124 -0.2901017  0.336168164
## Speed      -0.05175766  1.00000000 0.875571663  0.6458476  0.040389919
## Height      0.08198312  0.87557166 1.000000000  0.5504160  0.006564424
## Length     -0.29010167  0.64584765 0.550415985  1.0000000 -0.187121911
## Inversions  0.33616816  0.04038992 0.006564424 -0.1871219  1.000000000

There appears to be moderate multicollinearity between some of my variables, namely between Speed and Height with a value of 0.876, Speed and Length with a value of 0.646, and Length and Height with a value of 0.550. These values should not be severe enough to interfere with using Lasso regression, however. It is good at handling moderate collinearity.

y<- coasters$Speed
x<- data.matrix(coasters[, c('Height', 'Length', 'Track', 'Inversions')]) 

Lasso regression requires data to be input in the form of a matrix, allowing for efficient computation and uniform data types.

fit<- glmnet(x, y)
plot(fit)

This plot shows the coefficient paths for my Lasso regression. Each line represents a predictor in my dataset, and the L1 Norm represents the sum of the absolute values of the coefficients as the regularization penalty decreases from left to right. The black and red lines remain close to zero across the entire range of L1 Norm values, indicating they are less influential predictors. The blue and green lines have steeper slopes, indicating that the predictors and their coefficients are likely to be more influential.

To begin building my model, I randomly split my dataset into a training group containing 80% of my observations, and a testing group made up of the remaining 20%. I have set the seed to allow my results to be reproduced.

set.seed(76)
train_index <- sample(1:nrow(coasters), size = 0.8 * nrow(coasters))
train_data <- coasters[train_index, ]
test_data <- coasters[-train_index, ]
dim(train_data)
## [1] 192   7
dim(test_data)
## [1] 48  7

The training group consists of 192 roller coasters from my dataset, and the testing group consists of the remaining 48 coasters.

x_train <- model.matrix(Speed ~ Length + Height + Track + Inversions, data = train_data)[,-1]
y_train <- train_data$Speed
x_test <- model.matrix(Speed ~ Length + Height + Track + Inversions, data = test_data)[,-1]
y_test <- test_data$Speed

Now that my data has been prepared for Lasso regression, separated and arranged into matrix form, I can fit my regression model.

Now that the predictor variables have been arranged as columns in the input matrix, the glmnet package is able to standardize them all to have a zero mean and standard deviation of one. Unless specified otherwise, this standardization is performed by default before solving the Lasso regression optimization problem.

lasso_model <- glmnet(x_train, y_train, alpha = 1)
cv_lasso_model <- cv.glmnet(x_train, y_train, alpha = 1)

The training data is used to fit the model, and ten-fold cross validation is performed in order to find the optimal value of \(\lambda\). In this ten-fold cross validation, the training data is split into ten equal parts with the model built in 80% of the data and then tested in the remaining 20%. Each of the ten folds serves as a test set, and the mean square error (MSE) is averaged across all folds. This process is performed to further mitigate overfitting of the model.

best_lambda_min <- cv_lasso_model$lambda.min
cat("The value of lambda that minimizes MSE is", best_lambda_min, ".\n")
## The value of lambda that minimizes MSE is 0.05446444 .
best_lambda_1se <- cv_lasso_model$lambda.1se
cat("The value of lambda within one standard error of the minimum MSE is", best_lambda_1se, ".")
## The value of lambda within one standard error of the minimum MSE is 4.737039 .

Performing this ten-fold cross validation produces two relevant values of \(\lambda\), the value that minimizes MSE and the largest value within one standard error of the minimum MSE. These values are plotted below.

plot(cv_lasso_model)

The first vertical dotted line on the left represents the value of \(\lambda\) that minimizes MSE, the second line on the right represents the largest value of \(\lambda\) within one standard error of the minimum MSE. The red dots represent the cross-validated MSE for specific values of \(\lambda\), and the grey error bars represent the variability of the MSE across all ten folds. The choice between the two values of \(\lambda\) is dependent on the goal of the model. For the most accurate predictions, the value that minimizes MSE is the best choice but it can also allow for more complex models. For a simple model that can be generalized more broadly, the largest \(\lambda\) within one standard error of the minimum MSE is ideal. It is the larger value and able to apply more scrutiny to the predictors, eliminating more weak predictors and generating a model that is simple and more easily interpreted. My dataset does not have a particularly large quantity of predictors, but also does not have a large quantity of observations. I would like to avoid overfitting, but my main goal is accuracy of predictions. Because of this, I have selected the value of \(\lambda\) that minimizes MSE. This may lead to my model being less easily generalized, but my data does not come from a random sample and I am not looking to make generalizations to significantly larger populations.

final_lasso_model <- glmnet(x_train, y_train, alpha = 1, lambda = best_lambda_min)
coef(final_lasso_model)
## 5 x 1 sparse Matrix of class "dgCMatrix"
##                      s0
## (Intercept) 25.62116117
## Length       0.00245759
## Height       0.21350531
## Track       -5.53583690
## Inversions   3.84609341

After fitting the final model to the training group using the optimal value of \(\lambda\), that which minimizes MSE, I am left with the estimated coefficients. From the output we get the following fitted equation:

\[\widehat{Speed} = 25.62 + 0.21Height + 0.002Length - 5.54Track + 3.85Inversions\]

No predictors were eliminated from the model using the \(\lambda\) that minimizes MSE, indicating that they are all influential on the Speed of a roller coaster.

Next, the testing group is brought back in. The input data from the test group, Length, Height, Track, and Inversions, is used in the final model to make predictions about the response variable, Speed. The errors are calculated from the difference between the actual values of the response variable in the test group and the predicted values found with the model.

predictions <- predict(final_lasso_model, newx = x_test)
predictions <- as.vector(predictions)
comparison <- data.frame(Actual = y_test, Predicted = predictions)
comparison$Error <- comparison$Actual - comparison$Predicted
print(comparison)
##    Actual Predicted       Error
## 1    44.7  39.01066   5.6893410
## 2    45.0  46.34176  -1.3417623
## 3    80.0  79.06763   0.9323684
## 4    67.0  74.97261  -7.9726069
## 5    40.0  48.56109  -8.5610903
## 6    50.0  54.45617  -4.4561692
## 7   100.0 103.68241  -3.6824053
## 8    43.5  42.25727   1.2427296
## 9    67.1  66.10228   0.9977216
## 10   62.1  50.58444  11.5155590
## 11   46.6  52.70308  -6.1030823
## 12   25.7  37.75928 -12.0592794
## 13   26.7  30.29248  -3.5924806
## 14    4.5  22.43848 -17.9384841
## 15   83.0  63.36306  19.6369435
## 16   62.0  58.07811   3.9218862
## 17   82.0  62.11046  19.8895387
## 18   70.0  75.06203  -5.0620301
## 19   40.4  38.10648   2.2935171
## 20   47.9  45.87062   2.0293786
## 21   40.0  45.84253  -5.8425316
## 22   74.6  70.51064   4.0893638
## 23   95.0 108.02121 -13.0212132
## 24   67.0  79.01290 -12.0129023
## 25   64.8  66.22538  -1.4253842
## 26   47.0  51.10263  -4.1026325
## 27   29.1  33.79638  -4.6963825
## 28   83.3  85.92283  -2.6228287
## 29   64.9  58.89875   6.0012454
## 30   35.0  36.18912  -1.1891183
## 31   36.0  35.46609   0.5339117
## 32   65.0  61.20420   3.7957968
## 33   56.0  62.07721  -6.0772054
## 34   49.7  53.50448  -3.8044788
## 35   55.0  52.40896   2.5910407
## 36   65.0  76.42992 -11.4299208
## 37  100.0 111.72515 -11.7251502
## 38   50.1  52.97465  -2.8746517
## 39   77.0  77.76541  -0.7654140
## 40   66.3  64.08210   2.2179044
## 41   62.0  58.62253   3.3774667
## 42   71.5  67.86350   3.6365041
## 43   55.9  56.84403  -0.9440260
## 44   74.0  77.69936  -3.6993641
## 45   21.7  27.71132  -6.0113174
## 46   62.1  59.07010   3.0299043
## 47   53.0  57.36730  -4.3672975
## 48   47.9  45.87062   2.0293786

It can be seen that there is a broad range of values of errors, some less than 1.0 and others greater than 10.

Results, Conclusion, Summary, Presentation

After finalizing my model, I use the predicted and the real values of the test data to calculate a relevant statistic to evaluate the effectiveness of my model, the Root Mean Squared Error (RMSE).

mse <- mean((y_test - predictions)^2)
rmse <- sqrt(mse)
cat("Root Mean Squared Error (RMSE):", rmse, "\n")
## Root Mean Squared Error (RMSE): 7.362322

According to my calculations, the RMSE is equal to 7.36. The RMSE is expressed in the same units as the response variable, in this case miles per hour. This RMSE of 7.36 indicates that the average difference between predicted speeds and actual speeds is 7.36 MPH. This is a relatively low value, especially in terms of roller coasters which are machines designed to move at high speeds. This value indicates to me that my model makes reasonably accurate predictions.

The R-Squared value is not used to evaluate models found via Lasso regression, as the R-squared value tends to favor models with many predictors. As we have learned, the addition of predictors (even with no impact on the model) will always increase the R-squared value. As Lasso regression aims to create simple models and eliminate predictors, the R-squared statistic is not a good means of evaluating model effectiveness when using Lasso.

residuals <- y_test - predictions
plot(residuals, main="Residuals Plot", xlab="Index", ylab="Residuals", col="pink")
abline(h=0, col="skyblue1")

In the plot of residuals and fitted values, the residuals appear to have random variation above and below zero. There may be slightly more observations below zero, but there do not appear to be any patterns or clusters to the data. This shows that my model does not exhibit signs of bias, and agrees with the RMSE that it is well-fitting for the data.

qqnorm(residuals, main = "QQ Plot of Residuals", xlab = "Theoretical Quantiles", ylab = "Sample Quantiles", col="plum")
qqline(residuals, col = "seagreen1")

The QQ plot of residuals follows a positive linear trend, with many of the data points falling directly on the line of best fit. There are a small number of departures from the trend at more extreme theoretical quantities, but the departures are not severe. The data appears to have constant variance, as expected.

Having evaluated the fit of my model, next I use my final model in order to make predictions regarding the estimated speed of roller coasters with certain specifications if they were built in 2015. Predictions made by my model are not able to be broadly generalized, as my data is not from a random sample and only includes coasters built in the year 2015. My model cannot make predictions about coasters built in other years, but I am able to generalize to most roller coasters built in 2015 and industry trends that were prominent that year.

My first prediction is for a hypothetical coaster built in 2015 with a steel track that is 5076 feet long, reaches a height of 415 feet, and features an inversion. My second prediction is for a hypothetical coaster built in 2015 with a wooden track that is 600 feet long, reaches a height of 130 feet, and does not feature an inversion.

new_data <- data.frame(
  Length = 5076,     
  Height = 415,     
  Track = 1,        
  Inversions = 1    
)
newdata <- as.matrix(new_data)
new_data2 <- data.frame(
  Length = 600,     
  Height = 130,     
  Track = 0,        
  Inversions = 0    
)
newdata2 <- as.matrix(new_data2)
predicted_speed <- predict(final_lasso_model, s = best_lambda_min, newx = newdata)
predicted_speed2 <- predict(final_lasso_model, s = best_lambda_min, newx = newdata2)
cat("The predicted speed for a new roller coaster with length 5076 feet, height 415 feet, a steel track, and one inversion is", predicted_speed, "mph.\n")
## The predicted speed for a new roller coaster with length 5076 feet, height 415 feet, a steel track, and one inversion is 125.0108 mph.
cat("The predicted speed for a new roller coaster with length 600 feet, height 130 feet, a wooden track, and no inversion is", predicted_speed2, "mph.\n")
## The predicted speed for a new roller coaster with length 600 feet, height 130 feet, a wooden track, and no inversion is 54.85141 mph.

According to my final model, the predicted speed of my first hypothetical coaster with the longer, taller steel track and inversion is 125.01 MPH. The predicted speed of my second hypothetical coaster with the shorter, lower wooden track and no inversion is 54.85 MPH.

Discussions, Critique

What did you learn about the problem or question you set out to investigate? What were the weaknesses and strengths of your analysis and this method?

In my analysis, I found that the Lasso regression model identified all included predictors, Height, Length, Track, and Inversions, contribute significantly to the model and have influence on the speed of roller coasters in 2015. My final model equation is:

\[\widehat{Speed}=25.62+0.21*Height+0.002*Length-5.54*Track+3.85*Inversions\]

From this final fitted equation, we can see that, on average, roller coaster speed increases by 0.21 MPH for every additional foot in height, and increases marginally by 0.002 MPH for every additional foot in length that its track is built. Coasters with an inversion are expected to be 3.85 MPH faster on average than coasters without an inversion. We can also see that coasters with steel tracks are estimated to have speeds that are 5.54 MPH slower than coasters with wooden tracks, holding all other factors constant. This result does not seem logical, I would expect steel roller coasters to be faster on average. This result may indicate the presence of confounding factors, and could warrant further investigation in the future.

The Root Mean Square Error (RMSE) of the model was calculated to be 7.36, indicating predicted roller coaster speeds are off by an average of 7.36 MPH. This indicates a relatively low average prediction error considering the nature of roller coasters and the high speeds that they reach. The residual and QQ plots do not depict any significant patterns or biases, further indicating the model to be well-fitting to the data.

My model predicts that a steel coaster with a length of 5076 feet, height of 415 feet, and one inversion is predicted to have a speed of 125.01 MPH, and that a wooden coaster with a length of 600 feet, height of 130 feet, and no inversions is predicted to have a speed of 54.85 MPH.

My analysis confirmed a positive relationship between roller coaster Speed and variables Height, Length, and Inversions. These results are logical and align with my expectations that longer, taller roller coasters tend to reach higher speeds and that coasters featuring an inversion likely face a higher minimum speed threshold to ensure safety. My analysis also revealed that coasters with steel tracks are associated with speeds that are lower on average. This data does not align with my expectations, and could reflect the effects of omitted variables or other nuances in my dataset and regarding roller coaster trends in 2015.

Lasso regression was able to effectively handle the moderate levels of multicollinearity among predictors in my dataset. No predictors were eliminated, but Lasso was able to shrink variable coefficients according to how predictive they are and determine the best model for my dataset. The model was trained on a large subset of my data, and after using ten-fold cross validation to find the optimal value of \(\lambda\) the final model was constructed and tested on a small subset of my data that had been set aside. It was shown to have a low RMSE and well-behaved plots of residuals, suggesting that Lasso was successful in avoiding a problem of severe overfitting.

However, Lasso regression is limited as it only produces linear regression models. It does not apply any transformations to variables beyond the default standardization, and it does not consider any potential interactions between data or potential non-linear relationships between explanatory and response variables. It built the best multiple linear regression model for my dataset, but does not rule out the possibility of an even better model that could be built with interaction or nonlinear term(s).

My analysis here successfully determined how roller coaster speed in 2015 was influenced by height, length, track material, and whether or not a coaster’s track featured an inversion. The model performed well for my training and testing groups from my dataset, but its findings are limited to the context of roller coasters that were opened in 2015 and cannot be generalized to roller coasters across time. Future research could expand the dataset to multiple years, incorporate variables such as curvature or ride type, and investigate the unexpected relationship between track material and speed. Lasso could have also determined the contribution of Park and Country variables, but due to the complexity of those variables I chose to exclude them. Future analysis with the inclusion of these variables could provide further context for trends in roller coaster design in 2015 as it differs by country or geographic region. It would also be interesting to see an analysis of roller coaster trends at the end of 2025 compared directly to these observed trends from 2015 to see how roller coaster design has changed over ten years of high technological advancement.