Hot Hands Probability Lab

Author

Lauren Wismann

Hot Hand Probability Lab

Load Libraries

library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.1.4     ✔ readr     2.1.5
✔ forcats   1.0.0     ✔ stringr   1.5.1
✔ ggplot2   3.5.1     ✔ tibble    3.2.1
✔ lubridate 1.9.3     ✔ tidyr     1.3.1
✔ purrr     1.0.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(openintro)
Loading required package: airports
Loading required package: cherryblossom
Loading required package: usdata

Data

data("kobe_basket")
head(kobe_basket)
# A tibble: 6 × 6
  vs     game quarter time  description                                    shot 
  <fct> <int> <fct>   <fct> <fct>                                          <chr>
1 ORL       1 1       9:47  Kobe Bryant makes 4-foot two point shot        H    
2 ORL       1 1       9:07  Kobe Bryant misses jumper                      M    
3 ORL       1 1       8:11  Kobe Bryant misses 7-foot jumper               M    
4 ORL       1 1       7:41  Kobe Bryant makes 16-foot jumper (Derek Fishe… H    
5 ORL       1 1       7:03  Kobe Bryant makes driving layup                H    
6 ORL       1 1       6:01  Kobe Bryant misses jumper                      M    

Exercise 1

What does a streak length of 1 mean, i.e. how many hits and misses are in a streak of 1? What about a streak length of 0?

# A streak length of 1 means achieving 1 hit before a miss. A streak length of 0 means not achieving a hit before a miss, or that a miss has occurred directly after a hit, or a miss followed a miss. 
kobe_streak <- calc_streak(kobe_basket$shot)
summary(kobe_streak)
     length      
 Min.   :0.0000  
 1st Qu.:0.0000  
 Median :0.0000  
 Mean   :0.7632  
 3rd Qu.:1.0000  
 Max.   :4.0000  

Create a bar graph

ggplot(data = kobe_streak, aes(x = length)) +
  geom_bar()

Exercise 2

summary(kobe_streak)
     length      
 Min.   :0.0000  
 1st Qu.:0.0000  
 Median :0.0000  
 Mean   :0.7632  
 3rd Qu.:1.0000  
 Max.   :4.0000  
# From the data summary, Kobe's average streak length is 0.732. Given that the data is discrete and that a streak of 0 is considered a miss, his typical streak would be 1. His longest streak, the maximum, is 4. The bar graph is right-skewed and unimodal. 

Simulations in R

outcomes <- c("heads", "tails")
sample(outcomes, size = 1, replace = TRUE)
[1] "heads"
sim_fair_coin <- sample(outcomes, size = 100, replace = TRUE)
sim_fair_coin
  [1] "tails" "heads" "tails" "heads" "tails" "tails" "heads" "tails" "tails"
 [10] "heads" "heads" "tails" "tails" "tails" "tails" "tails" "tails" "tails"
 [19] "heads" "heads" "heads" "heads" "tails" "tails" "tails" "tails" "heads"
 [28] "tails" "heads" "tails" "tails" "tails" "heads" "heads" "tails" "tails"
 [37] "heads" "tails" "tails" "heads" "heads" "tails" "tails" "tails" "tails"
 [46] "heads" "heads" "heads" "heads" "tails" "tails" "heads" "heads" "tails"
 [55] "tails" "heads" "heads" "heads" "tails" "tails" "tails" "tails" "heads"
 [64] "tails" "tails" "heads" "tails" "heads" "tails" "tails" "heads" "heads"
 [73] "heads" "tails" "heads" "tails" "heads" "heads" "tails" "tails" "tails"
 [82] "tails" "heads" "tails" "tails" "heads" "tails" "tails" "heads" "heads"
 [91] "tails" "tails" "tails" "heads" "heads" "tails" "tails" "heads" "tails"
[100] "heads"
table(sim_fair_coin)
sim_fair_coin
heads tails 
   42    58 

Exercise 3

set.seed(777)
sim_unfair_coin <- sample(outcomes, size = 100, replace = TRUE, prob = c(0.2, 0.8))
sim_unfair_coin
  [1] "tails" "tails" "tails" "heads" "tails" "tails" "tails" "tails" "heads"
 [10] "tails" "tails" "tails" "tails" "tails" "heads" "tails" "tails" "heads"
 [19] "tails" "tails" "tails" "tails" "tails" "tails" "tails" "tails" "tails"
 [28] "tails" "heads" "tails" "heads" "tails" "heads" "tails" "tails" "tails"
 [37] "tails" "tails" "tails" "tails" "tails" "tails" "tails" "tails" "tails"
 [46] "heads" "heads" "heads" "heads" "tails" "tails" "tails" "tails" "tails"
 [55] "tails" "tails" "tails" "heads" "tails" "tails" "tails" "tails" "heads"
 [64] "tails" "heads" "tails" "tails" "tails" "heads" "tails" "tails" "tails"
 [73] "heads" "tails" "tails" "heads" "tails" "tails" "tails" "tails" "tails"
 [82] "tails" "tails" "tails" "tails" "tails" "tails" "tails" "tails" "heads"
 [91] "tails" "tails" "tails" "heads" "tails" "heads" "tails" "tails" "heads"
[100] "heads"
table(sim_unfair_coin)
sim_unfair_coin
heads tails 
   22    78 
# After flipping the unfair coin 100 times, heads came up 22 times. 

Simulating the Independent Shooter

outcomes <- c("H", "M")
sim_basket <- sample(outcomes, size = 1, replace = TRUE)

Exercise 4

set.seed(666)
shot_outcomes <- c("H", "M")
sim_basket <- sample(shot_outcomes, size = 133, replace = TRUE, prob = c(0.45, 0.55))
table(sim_basket)
sim_basket
 H  M 
67 66 

Exercise 5

sim_streak <- calc_streak(sim_basket)
summary(sim_streak)
     length 
 Min.   :0  
 1st Qu.:0  
 Median :1  
 Mean   :1  
 3rd Qu.:2  
 Max.   :4  

Exercise 6

ggplot(data = sim_streak, aes(x = length)) +
  geom_bar()

# For the simulated shooter, their typical streak length is 1 and their longest streak length is 4. In this case, the bar graph in unimodal and right-skewed. 

Exercise 7

# Given that the probability and sample size would remain the same, I would expect the distribution to be similar. It would still be right-skewed and have similar spread. The sample size seems large enough and the shots are independent from one another, so I wouldn't expect any drastic differences due to random chance, but there will be some small amount of variability.

Exercise 8

# The distributions are similar to one another, and the longest streak lengths are are the same for the simulated shooter and Kobe. Given the similarities between the two and the fact that the shots by the simulated shooter are independent, the data does not suggest that the hot hands model fits Kobe's shooting pattern in the 2009 NBA Finals. Furthermore, this suggests that Kobe's shots are are independent from one another, meaning that making a shot did not increase the probability of him making the next.