Workshop 3, Stats for AI - Bloque 2

Time series components and the Random Walk model

Author

Alberto Dorantes D., Ph.D.

Published

October 1, 2026

Abstract
In this workshop we learn what a time series is, how to decompose it into trend, seasonality and an irregular component, and how to model, estimate, simulate and forecast a random walk with drift using the S&P 500 index.

1 Learning objectives

By the end of this session you should be able to:

  1. Identify the components of a time series (trend, seasonality, cycles and irregular movements) and separate them with a decomposition.
  2. Write the random walk model, and explain which part of it is predictable and which part is not.
  3. Estimate the parameters of a random walk with drift correctly, using returns.
  4. Run a Monte Carlo simulation with many paths, and use it to build a forecast with an uncertainty band.

Why does this matter for AI? A lot of data used in AI is ordered in time: sales, prices, sensor readings, web traffic, energy demand. Any model that predicts the next value of such a series (a regression, a neural network, an LLM-based agent) must beat a very simple benchmark: the random walk (also called the naive forecast). If it can’t, it adds nothing. Today we learn that benchmark.

2 What is a time series?

A time series is a variable that measures an attribute of one subject over time, at regular intervals. Examples: the daily close of the S&P 500 since 2008, the monthly tons sold of a cereal brand, the quarterly GDP of Mexico.

The order of the observations matters. In a cross-sectional dataset you can shuffle the rows without losing information. In a time series, if you shuffle the rows you destroy the information. That has consequences for AI. For example, you cannot do a random train/test split with time-series data; we will come back to this in Workshop 5.

2.1 Components of a time series

We usually describe a time series with four components:

  1. Trend: a persistent upward or downward movement.
  2. Seasonality: systematic patterns that repeat every year (or every week, every day), such as higher candy sales in October–December.
  3. Cycles: up-and-down movements longer than one year that do not have a fixed period, such as business cycles (expansions and recessions).
  4. Irregular (or remainder): idiosyncratic movements that we cannot predict. We model this component as a random shock with mean zero and a certain standard deviation.

Trend and seasonality are usually the easiest components to model. Cycles are much harder to predict because their length and depth change from one cycle to another. The irregular component is, by definition, unpredictable.

2.2 Decomposing a real series

Let’s look at a real-world example. We download a monthly dataset with the volume sales (in tons) of several brands of one consumer-product category:

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import requests
from io import BytesIO

url = "http://www.apradie.com/ec2004/salesbrands3.xlsx"
# Some servers reject requests that do not look like a browser:
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36'}
response = requests.get(url, headers=headers)
response.raise_for_status()
sales_brand = pd.read_excel(BytesIO(response.content), sheet_name='Resultado')

We convert the date column to a date type and set it as the index of the data frame. With asfreq('MS') we tell pandas that this is a monthly series whose dates are the first day of each month (Month Start):

sales_brand['date'] = pd.to_datetime(sales_brand['date'])
sales_brand = sales_brand.set_index('date').sort_index().asfreq('MS')
sales_brand.head()
            qbrand1tons  salesbrand1  ...  tonstotal  salestotal
date                                  ...                       
2015-01-01       106.62     60780.72  ...    1148.41   302762.85
2015-02-01       105.82     60411.47  ...    1274.85   327076.83
2015-03-01        88.82     53174.89  ...    1085.59   298694.21
2015-04-01        86.69     50861.69  ...     961.22   267792.19
2015-05-01        87.23     51014.23  ...     997.47   266576.37

[5 rows x 14 columns]

Let’s plot the volume sales of brand 6:

plt.figure(figsize=(12, 5))
plt.plot(sales_brand.index, sales_brand['qbrand6tons'])
plt.title('Volume sales of brand 6 (tons)')
plt.xlabel('Date')
plt.ylabel('Tons')
plt.grid(True)
plt.show()

We can see a growing trend and a pattern that seems to repeat every year. To separate these components we use the STL decomposition (Seasonal-Trend decomposition using LOESS). We apply it to the natural logarithm of sales. With logs, a seasonal effect that is a percentage of the level (for example, “December sales are 20% above average”) becomes a constant additive effect, which is what STL assumes.

from statsmodels.tsa.seasonal import STL

log_q6 = np.log(sales_brand['qbrand6tons'])
stl_q6 = STL(log_q6, period=12, robust=True).fit()

fig = stl_q6.plot()
fig.set_size_inches(12, 8)
plt.show()

The decomposition writes the log series as trend + seasonal + remainder:

  • The trend panel shows the long-term direction. Cycles, if any, also show up here as slow waves around the trend.
  • The seasonal panel shows the pattern that repeats every 12 months.
  • The remainder (residual) panel is the irregular component: what is left after removing trend and seasonality.

How strong are the trend and the seasonality? A common way to measure it (Hyndman and Athanasopoulos 2021) compares the variance of the remainder with the variance of each component plus the remainder:

F_{trend} = \max\left(0,\ 1-\frac{Var(R_t)}{Var(T_t+R_t)}\right) \qquad F_{seasonal} = \max\left(0,\ 1-\frac{Var(R_t)}{Var(S_t+R_t)}\right)

Values close to 1 indicate a strong component; values close to 0 indicate a weak component.

R = stl_q6.resid
F_trend = max(0, 1 - R.var() / (stl_q6.trend + R).var())
F_seas = max(0, 1 - R.var() / (stl_q6.seasonal + R).var())
print(f"Strength of trend:       {F_trend:.3f}")
Strength of trend:       0.523
print(f"Strength of seasonality: {F_seas:.3f}")
Strength of seasonality: 0.875

For brand 6 the strength of the trend is 0.523 and the strength of the seasonality is 0.875.

The irregular component is what makes any series partially unpredictable. The simplest model where everything except a constant drift is irregular is the random walk, which we study next.

3 The Random Walk model

3.1 The model

The random walk hypothesis in Finance (Fama 1965) states that the natural logarithm of a stock price behaves like a random walk with drift. Let Y_t be the log price at day t:

Y_t = \phi_0 + Y_{t-1} + \varepsilon_t, \qquad \varepsilon_t \sim N(0, \sigma_\varepsilon^2) \text{ independent over time}

  • \varepsilon_t is the shock of day t: the net effect on the log price of all the news of that day (news about the firm, the industry, the economy…).
  • \phi_0 is the drift. If \phi_0>0 the series tends to grow; if \phi_0<0 it tends to decline; if \phi_0=0 we have a random walk without drift.
  • \sigma_\varepsilon is the volatility of the shock.

3.2 What is predictable and what is not?

It is common to say “a random walk cannot be predicted”. That is only half true. Re-arrange the equation by moving Y_{t-1} to the left:

\underbrace{Y_t - Y_{t-1}}_{r_t \text{ (log return)}} = \phi_0 + \varepsilon_t

The difference of log prices is the daily continuously compounded return, r_t. The equation says that the daily return is a constant plus a pure random shock. Then:

  • The change (the return) is unpredictable beyond its mean \phi_0. Past returns do not help to predict tomorrow’s return.
  • The level does have a best forecast: tomorrow’s expected log price is today’s log price plus the drift, E[Y_{t+1} \mid Y_t] = Y_t + \phi_0. That is why the random walk is also called the naive forecast.
  • What we cannot predict is how far from that forecast the real value will be. We will see that this uncertainty grows with the forecast horizon.

3.3 The value of a random walk after N periods

Let’s write the first values of the series starting from an initial value Y_0:

Y_1 = \phi_0 + Y_0 + \varepsilon_1

Y_2 = \phi_0 + Y_1 + \varepsilon_2 = 2\phi_0 + Y_0 + \varepsilon_1 + \varepsilon_2

If we continue until period N:

Y_N = Y_0 + N\phi_0 + \sum_{t=1}^{N}\varepsilon_t

The value after N periods equals the initial value, plus N times the drift, plus the sum of all the shocks. This equation tells us two important things. Imagine we could re-run history many times starting from the same Y_0 (each run is one possible path):

  • Since each shock has mean zero, the expected value across all possible paths is E[Y_N \mid Y_0] = Y_0 + N\phi_0.
  • Since the shocks are independent and have the same variance, the variance across all possible paths is

Var(Y_N \mid Y_0) = \sum_{t=1}^{N}Var(\varepsilon_t) = N\sigma_\varepsilon^2 \quad\Rightarrow\quad SD(Y_N \mid Y_0) = \sqrt{N}\,\sigma_\varepsilon

This is the famous square-root-of-time rule: the uncertainty about the future value grows with \sqrt{N}. The uncertainty about the value one year ahead (about 252 trading days) is \sqrt{252}\approx 15.9 times the uncertainty about tomorrow, not 252 times.

ImportantCareful: “variance across paths” is not “variance over time”

Var(Y_N)=N\sigma_\varepsilon^2 is the variance of Y_N across many possible histories. We only observe one history of the S&P 500. The standard deviation of the observed log prices over time is a different quantity. Because a random walk is not stationary, that standard deviation does not converge to anything meaningful.

3.4 Estimating the parameters correctly

The returns equation, r_t = \phi_0 + \varepsilon_t, gives the estimation recipe directly:

  • \hat\phi_0 = mean of the daily log returns
  • \hat\sigma_\varepsilon = standard deviation of the daily log returns

If we have N prices, we have N-1 returns. Notice that the mean of the returns telescopes, because the intermediate terms cancel:

\hat\phi_0=\frac{1}{N-1}\sum_{t=2}^{N}(Y_t-Y_{t-1})=\frac{Y_N-Y_1}{N-1}

So the drift is simply (last log price - first log price) / (number of returns).

4 The S&P 500 as a random walk

4.1 Data

We download the daily S&P 500 index from Yahoo Finance from 2008 to the end of 2025:

import yfinance as yf

sp500 = yf.download("^GSPC", start="2008-01-01", end="2026-01-01",
                    auto_adjust=True, progress=False)['Close'].squeeze().dropna()
sp500.name = 'SP500'
sp500.tail()
Date
2025-12-24    6932.049805
2025-12-26    6929.939941
2025-12-29    6905.740234
2025-12-30    6896.240234
2025-12-31    6845.500000
Name: SP500, dtype: float64

We will use the data in two parts, like a real forecasting exercise:

  • Training period (2008–2024): to estimate the model.
  • Test period (2025): to check how the forecast performs on data the model never saw.
lnsp = np.log(sp500)
ln_train = lnsp[:'2024-12-31']
ln_test = lnsp['2025-01-01':]

r_train = ln_train.diff().dropna()   # daily log returns

plt.figure(figsize=(12, 4))
plt.plot(r_train.index, r_train)
plt.title('Daily log returns of the S&P 500 (training period)')
plt.grid(True)
plt.show()

The log price clearly has a trend, but the returns fluctuate around a constant mean close to zero. We also see volatility clustering (2008–2009, 2020): calm periods and turbulent periods. The basic random walk assumes a constant \sigma_\varepsilon, so it does not capture this. That is a known limitation (models such as GARCH address it).

4.2 Are returns predictable from past returns?

If the random walk hypothesis holds, today’s return should not be correlated with yesterday’s return. The correlation of a series with its own past value is called autocorrelation (we will study it in detail in Workshops 4 and 5):

ac1 = float(r_train.autocorr(lag=1))
ac1
-0.12172626807671814

The lag-1 autocorrelation of daily returns is -0.122, very close to zero. Yesterday’s return tells us practically nothing about today’s return, which is consistent with the random walk hypothesis.

4.3 Estimating the parameters

phi0 = float(r_train.mean())
sigma = float(r_train.std())
y0 = float(ln_train.iloc[0])
N = len(ln_train)

# Check: the mean of returns equals (last - first)/(number of returns)
phi0_check = float((ln_train.iloc[-1] - ln_train.iloc[0]) / (N - 1))
print(f"phi0  = {phi0:.6f}   (check: {phi0_check:.6f})")
phi0  = 0.000328   (check: 0.000328)
print(f"sigma = {sigma:.6f}")
sigma = 0.012710

The estimated daily drift is \hat\phi_0= 0.0003, and the daily volatility of the shock is \hat\sigma_\varepsilon= 0.0127. In annual terms (252 trading days) that is a drift of about 8.3% and a volatility of about 20.2% (using the square-root-of-time rule).

4.4 Simulating one path

We simulate one random walk starting from the first log value of the training period. First with a for loop that follows the equation literally:

rng = np.random.default_rng(1)
shocks = rng.normal(0, sigma, N)

rw1 = np.zeros(N)
rw1[0] = y0
for t in range(1, N):
    rw1[t] = phi0 + rw1[t-1] + shocks[t]

The same can be done without a loop. Since Y_t = Y_0 + t\phi_0 + \sum_{s=1}^{t}\varepsilon_s, we use the cumulative sum (np.cumsum) of drift plus shocks:

steps = phi0 + shocks
steps[0] = 0                          # no shock on day 0
rw1_vec = y0 + np.cumsum(steps)
print(np.allclose(rw1, rw1_vec))      # both methods give the same path
True

We also simulate a random walk without drift (\phi_0=0) with the same shocks, to see the role of the drift:

steps0 = shocks.copy()
steps0[0] = 0
rw2 = y0 + np.cumsum(steps0)

plt.figure(figsize=(12, 5))
plt.plot(ln_train.index, ln_train, color='black', label='Real log S&P 500')
plt.plot(ln_train.index, rw1, label='Random walk with drift (rw1)')
plt.plot(ln_train.index, rw2, label='Random walk without drift (rw2)')
plt.title('One simulated path vs. the real log S&P 500')
plt.legend()
plt.grid(True)
plt.show()

One path alone does not tell us much: if you change the seed, the path changes completely. To learn what the model really says we need many paths. That is a Monte Carlo simulation.

4.5 Monte Carlo simulation: many paths

We simulate 1,000 paths for the whole training period. Each row of the matrix is one possible history:

n_paths = 1000
rng = np.random.default_rng(123)
step_matrix = phi0 + rng.normal(0, sigma, size=(n_paths, N))
step_matrix[:, 0] = 0                          # day 0 = y0 for every path
paths = y0 + np.cumsum(step_matrix, axis=1)    # one path per row

For each day we compute the 5th, 50th and 95th percentiles across the 1,000 paths. That gives a fan chart: the band where the model expects the series to be 90% of the time.

p05, p50, p95 = np.percentile(paths, [5, 50, 95], axis=0)

plt.figure(figsize=(12, 5))
for i in range(20):
    plt.plot(ln_train.index, paths[i], color='grey', alpha=0.25, linewidth=0.7)
plt.fill_between(ln_train.index, p05, p95, color='tab:blue', alpha=0.2, label='90% band of simulated paths')
plt.plot(ln_train.index, p50, color='tab:blue', label='Median simulated path')
plt.plot(ln_train.index, ln_train, color='black', label='Real log S&P 500')
plt.title('Monte Carlo simulation: 1,000 random walk paths')
plt.legend(loc='upper left')
plt.grid(True)
plt.show()


inside = float(np.mean((ln_train.values >= p05) & (ln_train.values <= p95)))

The real S&P 500 stays inside the 90% band 92.6% of the days. The real index looks like one more path of the random walk.

4.6 The square-root-of-time rule, visually

The band gets wider over time. Let’s compare the standard deviation across paths at each day with the theoretical value \sigma_\varepsilon\sqrt{t}:

t_idx = np.arange(N)
sd_across = paths.std(axis=0)

plt.figure(figsize=(12, 5))
plt.plot(t_idx, sd_across, label='SD across the 1,000 simulated paths')
plt.plot(t_idx, sigma * np.sqrt(t_idx), '--', label=r'Theory: $\sigma_\varepsilon\sqrt{t}$')
plt.xlabel('Days since the start')
plt.ylabel('Standard deviation of $Y_t$')
plt.legend()
plt.grid(True)
plt.show()

The simulated and theoretical curves overlap. This is Var(Y_N)=N\sigma_\varepsilon^2 in a picture.

5 Forecasting with the random walk

Now we use the model as a forecasting tool. Standing at the last day of 2024, what does the random walk say about 2025?

The analytical answer, for h days ahead, is:

Y_{T+h} \mid Y_T \sim N\left(Y_T + h\phi_0,\ h\sigma_\varepsilon^2\right)

We can also get it by Monte Carlo, which is the approach that still works for more complex models:

yT = float(ln_train.iloc[-1])
h = len(ln_test)

rng = np.random.default_rng(2025)
fc_paths = yT + np.cumsum(phi0 + rng.normal(0, sigma, size=(5000, h)), axis=1)
f05, f50, f95 = np.percentile(fc_paths, [5, 50, 95], axis=0)

plt.figure(figsize=(12, 5))
plt.plot(sp500['2024-01-01':'2024-12-31'], color='black', label='S&P 500 (2024)')
plt.plot(ln_test.index, np.exp(f50), color='tab:blue', label='Median forecast')
plt.fill_between(ln_test.index, np.exp(f05), np.exp(f95), color='tab:blue', alpha=0.2, label='90% forecast band')
plt.plot(sp500['2025-01-01':], color='tab:red', label='Actual 2025')
plt.title('Random walk forecast for 2025 vs. what actually happened')
plt.legend(loc='upper left')
plt.grid(True)
plt.show()

Let’s answer some concrete questions with the simulation and compare with the analytical formula:

from scipy.stats import norm

end_sim = fc_paths[:, -1]
actual_end = float(ln_test.iloc[-1])

p_up_mc = float(np.mean(end_sim > yT))
p_up_theory = float(norm.cdf(h * phi0 / (sigma * np.sqrt(h))))
band_low, band_high = np.exp(np.percentile(end_sim, [5, 95]))
pct_actual = float(np.mean(end_sim < actual_end))

print(f"Trading days in the forecast: {h}")
Trading days in the forecast: 250
print(f"P(S&P ends 2025 above its end-2024 level): MC = {p_up_mc:.3f}, theory = {p_up_theory:.3f}")
P(S&P ends 2025 above its end-2024 level): MC = 0.655, theory = 0.658
print(f"90% band for the end-2025 level: {band_low:,.0f} to {band_high:,.0f}")
90% band for the end-2025 level: 4,517 to 8,945
print(f"Actual end-2025 level: {np.exp(actual_end):,.0f}  (percentile within the simulation: {pct_actual:.2f})")
Actual end-2025 level: 6,845  (percentile within the simulation: 0.63)

Interpretation:

  • The model gives a probability of about 0.65 that the S&P 500 ends 2025 above its end-2024 level. It is above 0.5 because of the positive drift, but far from certain.
  • The 90% band for the end of 2025 is wide: from about 4,517 to 8,945 points. That width is the honest uncertainty of the model.
  • The actual end-2025 value fell at percentile 0.63 of the simulated distribution.

Discussion: a model that predicts the direction of the market correctly 55% of the time is not useless, but the random walk tells us how much uncertainty any prediction will have. Any AI model that claims to forecast prices must beat this naive benchmark out of sample.

6 A look ahead: what if the coefficient of Y_{t-1} is not 1?

The random walk gives a weight of exactly 1 to yesterday’s value. What happens if we write

Y_t = \phi_0 + \phi_1 Y_{t-1} + \varepsilon_t

with \phi_1 different from 1? This is the AR(1) model, and the value of \phi_1 changes the behavior of the series completely. It is the starting point of Workshop 4.

7 CHALLENGE

Work in your Notebook. Report your numbers and interpret them in your own words.

  1. Data and estimation. Choose one of the following: the Mexican stock index IPC (^MXX), Walmart de México (WALMEX.MX) or a US stock of your choice. Download daily data from 2015-01-01 to 2025-12-31. Using 2015–2024 as the training period:

    1. Estimate \phi_0 and \sigma_\varepsilon with daily log returns, and express both in annual terms.
    2. Compute the lag-1 autocorrelation of the daily returns. Is it consistent with the random walk hypothesis? Explain.
  2. Monte Carlo forecast. Simulate at least 2,000 paths for 2025 starting from the last log price of 2024.

    1. Plot the fan chart (5th, 50th and 95th percentiles in price scale) together with the actual 2025 prices.
    2. Report the probability that your asset ends 2025 above its end-2024 level, by Monte Carlo and with the analytical formula. Do they agree?
    3. At which percentile of your simulated distribution did the actual end-2025 price fall? Was 2025 an “unusual” year for your asset according to the model?
  3. Square-root-of-time. Using your simulation, compute the standard deviation of the simulated log price 21 trading days (about 1 month) ahead and 252 trading days (about 1 year) ahead. Compute their ratio and compare it with \sqrt{252/21}. Explain why the ratio is not 12.

  4. Reflection (max. 10 lines). Based on your results, what does the random walk imply for an AI model that tries to predict tomorrow’s price of your asset? What would that model need to show to be useful?

8 References

Fama, Eugene F. 1965. “The Behavior of Stock-Market Prices.” The Journal of Business 38 (1): 34–105.
Hyndman, Rob J., and George Athanasopoulos. 2021. Forecasting: Principles and Practice. 3rd ed. OTexts. https://otexts.com/fpp3/.