Question 7.1

Describe a situation or problem from your job, everyday life, current events, etc., for which exponential smoothing would be appropriate. What data would you need? Would you expect the value of  (the first smoothing parameter) to be closer to 0 or 1, and why?


Answer:

In my previous job as an installation engineer at a solar energy company, we installed solar panels and inverters for houses and mid-sized businesses. A big challenge we often faced was figuring out how much solar energy each place would need. Since every house is different, even those in the same area might need different amounts of resources.

To get a handle on this, we used to look at a bunch of different factors:

  • How much energy they produced each day
  • The usual sunlight and wind speed
  • Where they were located
  • How old and in what condition their equipment was
  • How their energy production changed with the seasons


Exponential smoothing could be highly effective for forecasting daily energy production within this context. This method would utilize historical data on daily energy production, along with variables such as average sunlight, wind speed, and seasonal variations. Given the unpredictable nature of weather conditions which can cause significant fluctuations in energy production, I would recommend setting the smoothing parameter, α, closer to 0. A smaller α value helps in giving greater emphasis to older observations, thus smoothing out short-term volatility and providing a more consistent and stable forecast.

Question 7.2

Using the 20 years of daily high temperature data for Atlanta (July through October) from Question 6.2 (file temps.txt), build and use an exponential smoothing model to help make a judgment of whether the unofficial end of summer has gotten later over the 20 years. (Part of the point of this assignment is for you to think about how you might use exponential smoothing to answer this question. Feel free to combine it with other models if you’d like to. There’s certainly more than one reasonable approach.)

Note: in R, you can use either HoltWinters (simpler to use) or the smooth package’s es function (harder to use, but more general). If you use es, the Holt-Winters model uses model=”AAM” in the function call (the first and second constants are used “A”dditively, and the third (seasonality) is used “M”ultiplicatively; the documentation doesn’t make that clear).

Answer:

The goal is to employ an Exponential Smoothing model to refine the daily temperature data by minimizing random fluctuations, followed by applying a change detection technique on this processed data. This will help us determine if there has been a shift in the timing of summer’s end over a 20-year period.

The process can be broken down into three key steps:

  • Step 1: Data Preparation - Organize and set up the data suitable for building an Exponential Smoothing model.
  • Step 2: Data Smoothing - Apply the Exponential Smoothing model to the prepared data to filter out the noise and enhance the underlying trends.
  • Step 3: Change Detection - Conduct change detection analysis on the smoothed data to explore if there’s a noticeable trend indicating a later end to summer across the years. (This will be in the form of CUSUM)
# Load data
temps_df <- read.table("temps.txt",stringsAsFactors = FALSE, header = TRUE)

# Needs to be converted as time series as the data in the df is not suitable to work with
temps_vec <- as.vector(unlist(temps_df[,2:21]))
head(temps_vec)
## [1] 98 97 97 90 89 93
plot(temps_vec)

temps_ts <- ts(temps_vec, start=1996,frequency = 123)
head(temps_ts)
## [1] 98 97 97 90 89 93
plot(temps_ts)

temps_hw <- HoltWinters(temps_ts, alpha = NULL, beta = NULL, gamma = NULL, seasonal = "multiplicative")
summary(temps_hw)
##              Length Class  Mode     
## fitted       9348   mts    numeric  
## x            2460   ts     numeric  
## alpha           1   -none- numeric  
## beta            1   -none- numeric  
## gamma           1   -none- numeric  
## coefficients  125   -none- numeric  
## seasonal        1   -none- character
## SSE             1   -none- numeric  
## call            6   -none- call
plot(temps_hw)

head(temps_hw$fitted)
##          xhat    level        trend   season
## [1,] 87.23653 82.87739 -0.004362918 1.052653
## [2,] 90.42182 82.15059 -0.004362918 1.100742
## [3,] 92.99734 81.91055 -0.004362918 1.135413
## [4,] 90.94030 81.90763 -0.004362918 1.110338
## [5,] 83.99917 81.93634 -0.004362918 1.025231
## [6,] 84.04496 81.93247 -0.004362918 1.025838
tail(temps_hw$fitted)
##             xhat    level        trend    season
## [2332,] 76.54551 87.81303 -0.004362918 0.8717307
## [2333,] 69.70436 81.07435 -0.004362918 0.8598048
## [2334,] 57.02909 71.26750 -0.004362918 0.8002607
## [2335,] 72.14646 87.37935 -0.004362918 0.8257107
## [2336,] 73.89293 85.77627 -0.004362918 0.8615051
## [2337,] 75.83100 82.99285 -0.004362918 0.9137532
temps_hw_sf <- matrix(temps_hw$fitted[,4], nrow=123)
33
## [1] 33
head (temps_hw_sf)
##          [,1]     [,2]     [,3]     [,4]     [,5]     [,6]     [,7]     [,8]
## [1,] 1.052653 1.049468 1.120607 1.103336 1.118390 1.108172 1.140906 1.140574
## [2,] 1.100742 1.099653 1.108025 1.098323 1.110184 1.116213 1.126827 1.154074
## [3,] 1.135413 1.135420 1.139096 1.142831 1.143201 1.138495 1.129678 1.156092
## [4,] 1.110338 1.110492 1.117079 1.125774 1.134539 1.126117 1.130758 1.137722
## [5,] 1.025231 1.025233 1.044684 1.067291 1.084725 1.097239 1.115055 1.103877
## [6,] 1.025838 1.025722 1.028169 1.042340 1.053954 1.067494 1.080203 1.094312
##          [,9]    [,10]    [,11]    [,12]    [,13]    [,14]    [,15]    [,16]
## [1,] 1.125438 1.122063 1.161415 1.198102 1.198910 1.243012 1.243781 1.238435
## [2,] 1.142187 1.131889 1.144549 1.134661 1.153433 1.165431 1.172935 1.190735
## [3,] 1.165657 1.147982 1.149459 1.135756 1.153310 1.155197 1.157286 1.169773
## [4,] 1.150639 1.146992 1.142497 1.150162 1.151169 1.157751 1.163844 1.159343
## [5,] 1.120818 1.133733 1.132167 1.142714 1.139244 1.112909 1.132435 1.132045
## [6,] 1.102680 1.092178 1.075766 1.088547 1.082185 1.103092 1.115071 1.118575
##         [,17]    [,18]    [,19]
## [1,] 1.300204 1.290647 1.254521
## [2,] 1.191956 1.219190 1.228826
## [3,] 1.189915 1.172309 1.169045
## [4,] 1.166605 1.167993 1.158956
## [5,] 1.145230 1.168161 1.170449
## [6,] 1.121598 1.134962 1.145475
34
## [1] 34
temps_hw_smoothed <- matrix(temps_hw$fitted[,1], nrow=123)
35
## [1] 35
plot(temps_hw_smoothed)

Question 8.2

Describe a situation or problem from your job, everyday life, current events, etc., for which a linear regression model would be appropriate. List some (up to 5) predictors that you might use.

Answer:

One situation where a linear regression model would be appropriate is predicting monthly electricity consumption for a household. This can be useful for energy management and planning, as well as for budgeting purposes. Here are some predictors that could be used in this model:

Average Monthly Temperature: Higher or lower temperatures can increase energy usage due to heating or cooling requirements. Household Size: More occupants typically lead to higher energy consumption. Number of Electrical Appliances: The presence of more appliances can directly increase electricity usage. Square Footage of the House: Larger homes generally require more energy for heating, cooling, and lighting. Monthly Usage of Major Appliances: The amount of time major appliances like air conditioners, heaters, and washing machines are used can significantly affect electricity consumption.