Exercise 3.1

Consider the GDP information in global_economy. Plot the GDP per capita for each country over time. Which country has the highest GDP per capita? How has this changed over time?

global_economy |>
  autoplot(GDP / Population) +
  theme(legend.position = "none") +
  labs(title = "GDP per Capita by Country", x = "Year", y = "US$")
## Warning: Removed 3242 rows containing missing values or values outside the scale range
## (`geom_line()`).

With 263 series on one plot (Country includes World Bank regional/income aggregates alongside actual countries), individual lines are impossible to tell apart, but the shape is still informative: almost every line sits low near the bottom, with a handful climbing well above the rest.

global_economy |>
  as_tibble() |>
  filter(!is.na(GDP), !is.na(Population)) |>
  mutate(GDP_per_capita = GDP / Population) |>
  group_by(Country) |>
  slice_max(GDP_per_capita, n = 1) |>
  ungroup() |>
  arrange(desc(GDP_per_capita)) |>
  select(Country, Year, GDP_per_capita) |>
  head(5)
## # A tibble: 5 × 3
##   Country           Year GDP_per_capita
##   <fct>            <dbl>          <dbl>
## 1 Monaco            2014        185153.
## 2 Liechtenstein     2014        179308.
## 3 Luxembourg        2014        119225.
## 4 Norway            2013        103059.
## 5 Macao SAR, China  2014         94004.

Monaco has the highest GDP per capita of any country in any single year in this data, peaking near US$185,000 in 2014 (shown above). But “which country has the highest GDP per capita” changes depending on the year, so a single-country plot understates how this has actually changed over time. Tracking the yearly leader instead:

gdp_pc <- global_economy |> mutate(GDP_per_capita = GDP / Population)

leaders <- gdp_pc |>
  as_tibble() |>
  filter(is.finite(GDP_per_capita)) |>
  group_by(Year) |>
  slice_max(GDP_per_capita, n = 1, with_ties = FALSE) |>
  ungroup()

leaders |>
  mutate(previous = lag(Country, default = first(Country)),
         period = cumsum(Country != previous)) |>
  group_by(period, Country) |>
  summarise(From = min(Year), To = max(Year), .groups = "drop") |>
  select(-period)
## # A tibble: 11 × 3
##    Country               From    To
##    <fct>                <dbl> <dbl>
##  1 United States         1960  1964
##  2 Kuwait                1965  1966
##  3 United States         1967  1969
##  4 Monaco                1970  1975
##  5 United Arab Emirates  1976  1977
##  6 Monaco                1978  2012
##  7 Liechtenstein         2013  2013
##  8 Monaco                2014  2014
##  9 Liechtenstein         2015  2015
## 10 Monaco                2016  2016
## 11 Luxembourg            2017  2017
gdp_pc |>
  filter(Country %in% c("Monaco", "Liechtenstein", "Luxembourg")) |>
  autoplot(GDP_per_capita) +
  labs(title = "GDP per Capita: the Recurring Leaders", y = "US$")
## Warning: Removed 22 rows containing missing values or values outside the scale range
## (`geom_line()`).

gdp_pc |>
  as_tibble() |>
  filter(Year == 2017,
         Country %in% c("Monaco", "Liechtenstein", "Luxembourg")) |>
  select(Country, Year, GDP, GDP_per_capita)
## # A tibble: 3 × 4
##   Country        Year          GDP GDP_per_capita
##   <fct>         <dbl>        <dbl>          <dbl>
## 1 Liechtenstein  2017          NA             NA 
## 2 Luxembourg     2017 62404461275.        104103.
## 3 Monaco         2017          NA             NA

The leader has actually changed hands several times: the United States led in the early 1960s, Kuwait briefly overtook it in 1965-66, and then Monaco held the top spot almost continuously from 1970 through 2016 - interrupted only by the United Arab Emirates in 1976-77 and by Liechtenstein in 2013 and 2015. Monaco’s run wasn’t a smooth climb either - it includes real dips (after 2008, and again after the 2014 peak). The leader table stops at 2017, and in that last year Luxembourg shows the highest value - but the table above shows why: Monaco and Liechtenstein simply have no GDP figure recorded for 2017 (NA), while Luxembourg’s is present. That gap means the data can’t actually confirm or rule out a real change in ranking that year - it just can’t answer the question, so I wouldn’t read “Luxembourg leads in 2017” as a real changing-of-the-guard without another source to fill in what Monaco and Liechtenstein’s 2017 figures actually were.

Exercise 3.2

For each of the following series, make a graph of the data. If transforming seems appropriate, do so and describe the effect.

United States GDP from global_economy.

us_gdp <- global_economy |> filter(Country == "United States")
us_gdp |> autoplot(GDP) + labs(title = "US GDP", y = "US$")

lambda_gdp <- us_gdp |>
  features(GDP, features = guerrero) |>
  pull(lambda_guerrero)
us_gdp |> autoplot(box_cox(GDP, lambda_gdp)) +
  labs(title = paste("US GDP, Box-Cox lambda =", round(lambda_gdp, 2)))

US GDP is annual data with no seasonal component to stabilise, just a smooth upward curve that gets steeper over time. Guerrero picks lambda = 0.28, and the transform reduces the curvature, compressing the later, larger values onto a scale closer to the earlier ones.

Slaughter of Victorian “Bulls, bullocks and steers” in aus_livestock.

vic_bulls <- aus_livestock |>
  filter(State == "Victoria", Animal == "Bulls, bullocks and steers")
vic_bulls |>
  autoplot(Count) +
  labs(title = "Victoria Bulls/Bullocks/Steers Slaughtered")

lambda_bulls <- vic_bulls |>
  features(Count, features = guerrero) |>
  pull(lambda_guerrero)
vic_bulls |> autoplot(box_cox(Count, lambda_bulls)) +
  labs(title = paste("Vic Bulls/Bullocks/Steers, Box-Cox lambda =",
                      round(lambda_bulls, 2)))

This series has substantial fluctuations and real changes in level over time rather than a strong trend. Guerrero returns a lambda (-0.04) close to a log transform, and the transformed plot still looks noisy, since a transform stabilises variance rather than smoothing out fluctuations. I’d use this lambda to express changes on a more relative scale.

Victorian Electricity Demand from vic_elec.

vic_elec |>
  autoplot(Demand) +
  labs(title = "Victoria Half-Hourly Electricity Demand")

lambda_elec <- vic_elec |>
  features(Demand, features = guerrero) |>
  pull(lambda_guerrero)
vic_elec |> autoplot(box_cox(Demand, lambda_elec)) +
  labs(title = paste("Vic Electricity Demand, Box-Cox lambda =",
                      round(lambda_elec, 2)))

vic_elec is half-hourly data with several seasonal periods stacked on top of each other (daily, weekly, and yearly), so a single Box-Cox lambda estimated across the whole three-year (2012-2014) series is a blunt tool - it can’t target any one of those periods specifically. The transformed series (lambda = 0.1) looks similar to the original in overall shape at this scale, so I’d want to check each seasonal period separately before relying on this transform.

Gas production from aus_production.

aus_production |>
  autoplot(Gas) +
  labs(title = "Australian Quarterly Gas Production")

lambda_gas <- aus_production |>
  features(Gas, features = guerrero) |>
  pull(lambda_guerrero)
aus_production |> autoplot(box_cox(Gas, lambda_gas)) +
  labs(title = paste("Gas Production, Box-Cox lambda =",
                      round(lambda_gas, 2)))

This is the clearest case for transforming: the seasonal swing in the raw series visibly grows as the level of gas production grows through the 1970s-80s. Guerrero picks lambda = 0.11, and the transformed plot shows a seasonal amplitude that stays roughly constant across the whole series instead of widening - exactly the effect a Box-Cox transform is meant to have.

Exercise 3.3

Why is a Box-Cox transformation unhelpful for the canadian_gas data?

canadian_gas |> autoplot(Volume) + labs(title = "Canadian Gas Production")

lambda_cg <- canadian_gas |>
  features(Volume, features = guerrero) |>
  pull(lambda_guerrero)
canadian_gas |> autoplot(box_cox(Volume, lambda_cg)) +
  labs(title = paste("Canadian Gas, Box-Cox lambda =", round(lambda_cg, 2)))

Box-Cox transforms are useful more generally for stabilising variance or straightening a curved trend (as with US GDP in 3.2, which has no seasonality at all) - but for the specific job of stabilising seasonal variance, they only help when that seasonal swing is a function of the series’ level: bigger values, bigger seasonal swings. Here that relationship doesn’t hold:

canadian_gas |>
  as_tibble() |>
  mutate(year = year(Month)) |>
  group_by(year) |>
  summarise(n_months = n(), level = mean(Volume),
            seasonal_spread = max(Volume) - min(Volume)) |>
  filter(n_months == 12,
         year %in% c(1973, 1986, 2004))
## # A tibble: 3 × 4
##    year n_months level seasonal_spread
##   <dbl>    <int> <dbl>           <dbl>
## 1  1973       12  8.24            1.82
## 2  1986       12  8.81            3.74
## 3  2004       12 18.0             2.55

I restricted this to complete years - the data end in February 2005, so 2005 only has two months and its range isn’t comparable to a full year’s. Between 1973 and 1986 the yearly average level is essentially flat (8.24 to 8.81), but the within-year range more than doubles (1.82 to 3.74). Then between 1986 and 2004 the level more than doubles (8.81 to 18.0, a 104% increase), but the range actually shrinks over that same stretch (3.74 to 2.55, down 32%). That range isn’t a pure measure of seasonality - it also picks up trend and irregular movement within the year - but it’s still telling: the swing grows while the level is flat, then shrinks while the level more than doubles, on a schedule of its own that tracks neither direction consistently. No single value of lambda can straighten that out, and the transformed plot above still shows the same changing-amplitude pattern as the original.

Exercise 3.4

What Box-Cox transformation would you select for your retail data (from Exercise 7 in Section 2.10)?

Homework 1 didn’t include exercise 2.7, so I’m building that retail subset here first, the same way the book does it, before I can answer this one.

set.seed(219562)
myseries <- aus_retail |>
  filter(`Series ID` == sample(aus_retail$`Series ID`, 1))
myseries |> distinct(State, Industry)
## # A tibble: 1 × 2
##   State      Industry                                     
##   <chr>      <chr>                                        
## 1 Queensland Cafes, restaurants and takeaway food services

My random draw landed on Queensland Cafes, Restaurants and Takeaway Food Services turnover.

myseries |> autoplot(Turnover) +
  labs(title = "Queensland Cafes/Restaurants/Takeaway Turnover")

lambda_retail <- myseries |>
  features(Turnover, features = guerrero) |>
  pull(lambda_guerrero)
myseries |> autoplot(box_cox(Turnover, lambda_retail)) +
  labs(title = paste("Queensland Cafes Turnover, Box-Cox lambda =",
                      round(lambda_retail, 2)))

Like the gas production series, this one’s seasonal swing clearly widens as the level of turnover grows over the decades. Guerrero selects lambda = 0.14, close to a log transform, and the transformed series shows a much more even seasonal amplitude across the whole history instead of ballooning toward the end - this is a reasonable candidate transform for this series, though I’d still check it against whatever specific model I ended up fitting rather than assuming it’s automatically right for any constant-variance method.

Exercise 3.5

For the following series, find an appropriate Box-Cox transformation in order to stabilise the variance. Tobacco from aus_production, Economy class passengers between Melbourne and Sydney from ansett, and Pedestrian counts at Southern Cross Station from pedestrian.

# as_tibble() first: summarise() on a bare tsibble computes per-index-group
# (one row per quarter) instead of collapsing to a single summary row.
aus_production |> as_tibble() |>
  summarise(n_na = sum(is.na(Tobacco)),
            last_observed = max(Quarter[!is.na(Tobacco)]))
## # A tibble: 1 × 2
##    n_na last_observed
##   <int>         <qtr>
## 1    24       2004 Q2
# The last nonmissing Tobacco observation is 2004 Q2, even though
# aus_production continues through 2010 Q2 for other series - drop the
# trailing NAs so they don't silently show up as a gap in the plot or get
# flagged by finite-value checks. The data don't say why later values are
# missing, only that they are.
tobacco <- aus_production |>
  select(Quarter, Tobacco) |>
  filter(!is.na(Tobacco))
tobacco |> autoplot(Tobacco) + labs(title = "Australian Tobacco Production")

tobacco |> gg_season(Tobacco) + labs(title = "Tobacco Production by Quarter")

lambda_tob <- tobacco |>
  features(Tobacco, features = guerrero) |>
  pull(lambda_guerrero)
tobacco |> autoplot(box_cox(Tobacco, lambda_tob)) +
  labs(title = paste("Tobacco, Box-Cox lambda =", round(lambda_tob, 2)))

The season plot shows tobacco production has a real, fairly strong quarterly pattern - not obvious from the raw autoplot alone, where a long rise-and-decline in level dominates the view. But strong seasonality and a strong Box-Cox transform aren’t the same thing: guerrero’s chosen lambda (0.93) is close to 1 (no transform), suggesting the seasonal amplitude doesn’t obviously grow or shrink with the level here. So I’d leave this one close to untransformed.

mel_syd_econ <- ansett |> filter(Airports == "MEL-SYD", Class == "Economy")
mel_syd_econ |> autoplot(Passengers) +
  labs(title = "Melbourne-Sydney Economy Passengers")

lambda_ansett <- mel_syd_econ |>
  features(Passengers, features = guerrero) |>
  pull(lambda_guerrero)
mel_syd_econ |> autoplot(box_cox(Passengers, lambda_ansett)) +
  labs(title = paste("MEL-SYD Economy, Box-Cox lambda =",
                      round(lambda_ansett, 2)))

This series has a few weeks with passenger counts dropping to near zero, which dominate the visual range more than any level-dependent seasonal swing. Guerrero picks lambda = 2, a fairly strong transform, but the near-zero weeks still stand out as a sharp interruption after transforming - the transform doesn’t remove that disruption, since it’s addressing level-dependent variance generally, not a specific gap in service.

scs <- pedestrian |> filter(Sensor == "Southern Cross Station")
scs |>
  autoplot(Count) +
  labs(title = "Southern Cross Station Pedestrian Counts")

lambda_ped_raw <- scs |>
  features(Count, features = guerrero) |>
  pull(lambda_guerrero)
n_undefined <- sum(!is.finite(box_cox(scs$Count, lambda_ped_raw)))
cat("zero counts:", sum(scs$Count == 0),
    "| non-finite after Box-Cox:", n_undefined, "\n")
## zero counts: 159 | non-finite after Box-Cox: 159

This is a real problem, not just a stylistic one: scs has 159 hours at exactly zero, and guerrero’s unshifted estimate (-0.25) is negative - box_cox(0, lambda) for a negative lambda is undefined, so applying that lambda directly produces 159 non-finite values, one for every zero-count hour. Adding 1 to every count before estimating and transforming avoids the zero-at-a-negative-power problem entirely:

scs <- scs |> mutate(Count_plus_one = Count + 1)
lambda_ped <- scs |>
  features(Count_plus_one, features = guerrero) |>
  pull(lambda_guerrero)

scs |> autoplot(box_cox(Count_plus_one, lambda_ped)) +
  labs(title = paste("Southern Cross Station, Box-Cox(Count+1), lambda =",
                      round(lambda_ped, 2)))

stopifnot(all(is.finite(box_cox(scs$Count_plus_one, lambda_ped))))

# which hours account for the 159 zero counts?
scs |> as_tibble() |> filter(Count == 0) |>
  mutate(hour = hour(Date_Time)) |> count(hour)
## # A tibble: 7 × 2
##    hour     n
##   <int> <int>
## 1     1     5
## 2     2    49
## 3     3    76
## 4     4    26
## 5     8     1
## 6     9     1
## 7    14     1

With the offset, lambda = -0.26 and every transformed value is finite (checked above). This is technically a Box-Cox transform of Count + 1, not of Count itself, so the offset is a real modelling choice worth stating explicitly rather than a detail to gloss over. It compresses the tall daytime peaks down closer to the low-count hours - two full years of hourly data plotted at this scale are too dense to clearly read a day/night or weekday/weekend pattern by eye, so I won’t claim that pattern is “clearly visible” here without a zoomed-in window to actually back it up; what the transform does verifiably accomplish is compressing the range while keeping every observation finite.

The 159 zero counts also aren’t purely an “overnight” phenomenon - most (156) fall in hours 1-4, but a handful sit at hours 8, 9, and 14, which is daytime. This data alone can’t tell me why those particular daytime hours read zero - a sensor outage is plausible, but so is something else - so I’ll just flag them as unexplained rather than guess at a cause.

Exercise 3.7

Consider the last five years of the Gas data from aus_production.

gas <- tail(aus_production, 5 * 4) |> select(Gas)

Plot the time series. Can you identify seasonal fluctuations and/or a trend-cycle?

gas |> autoplot(Gas) + labs(title = "Australian Gas Production, Last 5 Years")

There’s a clear quarterly seasonal pattern - Q3 consistently highest, Q1 and Q4 lower - riding on top of a gentle upward trend across the five years.

Use classical_decomposition with type = multiplicative to calculate the trend-cycle and seasonal indices.

dcmp <- gas |>
  model(classical_decomposition(Gas, type = "multiplicative")) |>
  components()
dcmp |> autoplot()
## Warning: Removed 8 rows containing missing values or values outside the scale range
## (`geom_line()`).

Do the results support the graphical interpretation from part a?

Yes - the extracted trend line rises gradually across the five years, though not at a constant rate: it visibly flattens out around 2007 Q4-2008 Q3 (shown in the trend panel above) before resuming its climb. The seasonal component repeats the same Q3-high, Q1-low shape every year, matching what’s visible by eye in the raw plot.

Compute and plot the seasonally adjusted data.

dcmp |>
  ggplot(aes(x = Quarter)) +
  geom_line(aes(y = Gas, colour = "Raw")) +
  geom_line(aes(y = season_adjust, colour = "Seasonally Adjusted")) +
  labs(title = "Gas Production: Raw vs Seasonally Adjusted",
       y = "Gas", colour = NULL)

The seasonally adjusted line removes the sharp quarterly zigzag from the raw series, leaving a much smoother climb that closely tracks the underlying trend - exactly what’s left once the repeating Q3-high, Q1-low pattern is divided back out.

Change one observation to be an outlier (e.g., add 300 to one observation), and recompute the seasonally adjusted data. What is the effect of the outlier?

gas_mid <- gas
gas_mid$Gas[10] <- gas_mid$Gas[10] + 300

dcmp_mid <- gas_mid |>
  model(classical_decomposition(Gas, type = "multiplicative")) |>
  components()

dcmp_mid |>
  ggplot(aes(x = Quarter)) +
  geom_line(aes(y = Gas, colour = "Raw (with outlier)")) +
  geom_line(aes(y = season_adjust, colour = "Seasonally Adjusted")) +
  labs(title = "Outlier in the Middle (2007 Q4)", y = "Gas", colour = NULL)

# seasonal index per quarter, before vs after the outlier
orig_seas <- dcmp |> as_tibble() |>
  mutate(q = quarter(Quarter)) |> distinct(q, seasonal)
mid_seas <- dcmp_mid |> as_tibble() |>
  mutate(q = quarter(Quarter)) |> distinct(q, seasonal)

bind_rows(original = orig_seas, with_outlier = mid_seas, .id = "case") |>
  pivot_wider(names_from = case, values_from = seasonal) |>
  arrange(q)
## # A tibble: 4 × 3
##       q original with_outlier
##   <int>    <dbl>        <dbl>
## 1     1    0.875        0.821
## 2     2    1.07         0.998
## 3     3    1.13         1.06 
## 4     4    0.925        1.12

The +300 spike is still clearly visible in the seasonally adjusted series, not smoothed away - classical decomposition can’t tell an outlier apart from a real seasonal shift. The Q4 seasonal index (printed above) rises from about 0.93 to 1.12: that index averages the detrended ratios for Q4, and since trend is unavailable at the series’ first two quarters (one of which is a Q4), only 4 of the 5 Q4 observations produce a usable ratio - the corrupted 2007 Q4 value is one of those 4. The outlier also nudges the nearby trend estimates it falls inside, and because all four quarters’ seasonal indices are normalized together, the effect isn’t confined to Q4 either - seasonally adjusted values change across all four quarters, not just the one quarter that actually changed.

Does it make any difference if the outlier is near the end rather than in the middle of the time series?

gas_end <- gas
gas_end$Gas[19] <- gas_end$Gas[19] + 300

dcmp_end <- gas_end |>
  model(classical_decomposition(Gas, type = "multiplicative")) |>
  components()

dcmp_end |>
  select(Quarter, Gas, trend, seasonal, season_adjust) |>
  print(n = 20)
## # A tsibble: 20 x 5 [1Q]
##    Quarter   Gas trend seasonal season_adjust
##      <qtr> <dbl> <dbl>    <dbl>         <dbl>
##  1 2005 Q3   221   NA     1.11           199.
##  2 2005 Q4   180   NA     0.889          202.
##  3 2006 Q1   171  200.    0.897          191.
##  4 2006 Q2   224  204.    1.10           203.
##  5 2006 Q3   233  207     1.11           209.
##  6 2006 Q4   192  210.    0.889          216.
##  7 2007 Q1   187  213     0.897          208.
##  8 2007 Q2   234  216.    1.10           213.
##  9 2007 Q3   245  219.    1.11           220.
## 10 2007 Q4   205  219.    0.889          231.
## 11 2008 Q1   194  219.    0.897          216.
## 12 2008 Q2   229  219     1.10           208.
## 13 2008 Q3   249  219     1.11           224.
## 14 2008 Q4   203  220.    0.889          228.
## 15 2009 Q1   196  222.    0.897          218.
## 16 2009 Q2   238  223.    1.10           216.
## 17 2009 Q3   252  263.    1.11           226.
## 18 2009 Q4   210  301     0.889          236.
## 19 2010 Q1   505   NA     0.897          563.
## 20 2010 Q2   236   NA     1.10           214.
dcmp_end |>
  ggplot(aes(x = Quarter)) +
  geom_line(aes(y = Gas, colour = "Raw (with outlier)")) +
  geom_line(aes(y = season_adjust, colour = "Seasonally Adjusted")) +
  labs(title = "Outlier Near the End (2010 Q1)", y = "Gas", colour = NULL)

Placement makes a difference here: the endpoint case produces an adjusted spike of about 563, versus about 449 for the same-size outlier in the middle. Classical decomposition’s moving-average trend needs data on both sides of a point, so the trend is NA for the first two and last two quarters of any window (printed above). At 2010 Q1, trend is NA, so that observation can’t directly contribute a detrended ratio to Q1’s seasonal index the way the middle-case outlier did for Q4 - though it still affects the preceding quarters’ trend estimates indirectly, since its raw value falls inside their moving-average windows. season_adjust is always Gas / seasonal - because the Q1 seasonal index changes only modestly, the re-estimated seasonal divisor offsets less of the added spike. That said, the two experiments corrupt different quarters (Q1 vs Q4) with different baseline seasonal indices, so this one comparison doesn’t prove endpoint outliers are always worse - it shows the mechanism, not a general rule.

Exercise 3.8

Recall your retail time series data (from Exercise 7 in Section 2.10). Decompose the series using X-11. Does it reveal any outliers, or unusual features that you had not noticed previously?

Exercise 3.4 already showed the raw autoplot() of this series. Adding a subseries view here as a second “before” baseline, alongside that raw plot, before running X-11:

myseries |> gg_subseries(Turnover) +
  scale_x_yearmonth(date_breaks = "10 years", date_labels = "%Y") +
  labs(title = "Queensland Cafes Turnover by Calendar Month")

Both views show real structure worth noting on their own - the growth rate isn’t constant (a faster climb in some stretches than others, and something closer to a plateau/decline near the most recent years) - but neither one makes late-1988 specifically stand out; at this scale a two-month dip inside a decades-long series with plenty of other level changes isn’t something I’d flag just from eyeballing it. Now the X-11 decomposition:

x11_dcmp <- myseries |>
  model(x11 = feasts::X_13ARIMA_SEATS(Turnover ~ x11())) |>
  components()

x11_dcmp |> autoplot()

x11_dcmp |>
  as_tibble() |>
  mutate(irregular_dev = abs(irregular - 1)) |>
  arrange(desc(irregular_dev)) |>
  select(Month, Turnover, trend, seasonal, irregular) |>
  head(5)
## # A tibble: 5 × 5
##      Month Turnover trend seasonal irregular
##      <mth>    <dbl> <dbl>    <dbl>     <dbl>
## 1 1988 Nov     87.2 118.     1.01      0.727
## 2 1988 Dec    114.  116.     1.15      0.857
## 3 1987 Jun     72.8  84.7    0.946     0.908
## 4 1991 Feb    145.  151.     0.888     1.08 
## 5 2010 Jul    576.  511.     1.04      1.08

The X-11 irregular component tells a different story than the subseries view above: the two biggest departures from 1 are both in late 1988 (November at 0.73, December at 0.86), meaning turnover that year came in noticeably below what trend and season alone would predict, for two months running - a dip that isn’t obvious in the subseries plot, since it sits inside an otherwise-growing stretch and doesn’t shift either month’s overall pattern enough to be visible by eye. The X-11 irregular component is exactly built to catch this kind of thing: it’s what’s left after trend and a seasonal estimate that’s allowed to evolve slowly are both removed, so a two-month underperformance like this stands out numerically even when it doesn’t visually break the overall growth pattern. The remaining three months in that table are far milder by comparison - June 1987 (0.91), February 1991 (1.08), and July 2010 (1.08) all sit within about 9% of 1, closer to ordinary month-to-month noise than to the two clear standouts in late 1988.

Exercise 3.9

Figures 3.19 and 3.20 show the result of decomposing the number of persons in the civilian labour force in Australia each month from February 1978 to August 1995.

The figures are at https://otexts.com/fpp3/decomposition-exercises.html.

Write about 3-5 sentences describing the results of the decomposition. Pay particular attention to the scales of the graphs in making your interpretation.

The labour force size rose from around 6,500 (thousand) in early 1978 to roughly 9,000 by 1995, and the trend panel mostly tracks that climb - but not at a constant pace: it visibly flattens into a plateau around 1991-1993 before resuming its upward climb toward the end of the series. The seasonal component repeats a broadly similar shape every year on a scale of roughly ±100 - December is clearly the highest month, and January and August are both consistently the lowest - but Figure 3.20 shows this shape isn’t perfectly fixed either: March’s peak is large and sharp in the early years but shrinks noticeably by the 1990s, while August’s dip gets somewhat deeper over the same period. The remainder panel is small (mostly within about ±50) for most of the series, but around 1991 there’s a cluster of unusually large negative values, including a sharp dip down to nearly -400 - much larger than the typical fluctuation elsewhere in that panel - along with a smaller, but still distinctly larger-than-typical, negative episode around 1993.

Is the recession of 1991/1992 visible in the estimated components?

Yes - the decomposition shows a slowdown in the trend and an unusual cluster of large negative remainder values both coinciding with the recession period, in both the trend and the remainder, not just the remainder. That’s consistent with the recession’s effect on the labour force showing up on two different timescales - a slower change absorbed into the trend, and a faster one that the trend-cycle’s smoothing window couldn’t track, left over in the remainder - though the decomposition alone doesn’t prove that causal story; it only shows that both panels behave unusually over the same stretch. The 1991 dip is visible even in the raw value panel, but reading the remainder panel’s own scale is what makes its size relative to typical remainder fluctuations easier to see: that -400 value would be unremarkable next to the trend’s own scale (which spans 6,500-9,000), but it clearly stands out against the remainder panel’s normal ±50 band.