Global Baseline Estimate

Approach

In this assignment, I will use the movie ratings data I collected from five participants for the six selected movies. The dataset contains ratings on a scale from 1 to 5, along with missing values for movies that participants had not seen. I will use this data to create a recommendation system using the Global Baseline Estimate method. I will use my ratings data to calculate user and movie effects and predict ratings for movies that participants have not seen. The movie with the highest predicted rating will be used as the recommendation for each participant.

One challenge I anticipate is calculating the Global Baseline Estimates correctly when participants have rated different numbers of movies. I will use the available ratings to calculate the overall, participant, and movie averages before using the model to predict the missing ratings.

Data Source: Primary movie rating data collected from five participants from the previous Movie Ratings assignment.

Codebase

library(DBI)
## Warning: package 'DBI' was built under R version 4.4.3
library(RPostgres)
## Warning: package 'RPostgres' was built under R version 4.4.3
library(dplyr)
## Warning: package 'dplyr' was built under R version 4.4.3
## 
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
## 
##     filter, lag
## The following objects are masked from 'package:base':
## 
##     intersect, setdiff, setequal, union
con <- dbConnect(
  RPostgres::Postgres(),
  dbname = "movie_ratings",
  host = "localhost",
  port = 5432,
  user = "postgres",
  password = rstudioapi::askForPassword("Enter PostgreSQL password")
)
# I checked the tables stored in the database. 
dbListTables(con)
## [1] "movies"       "participants" "ratings"

I used a SQL query to join the three tables and load the results into an R dataframe.

ratings_df <- dbGetQuery(
  con,
  "
  SELECT
    p.name AS participant,
      m.title AS movie,
    r.rating
  FROM ratings AS r
  JOIN participants AS p
    ON r.user_id = p.user_id
  JOIN movies AS m
    ON r.movie_id = m.movie_id
ORDER BY p.user_id, m.movie_id;
"
)

ratings_df
##    participant                     movie rating
## 1     Person 1 Spider-Man: Brand New Day      5
## 2     Person 1                   Sinners      4
## 3     Person 1               The Odyssey     NA
## 4     Person 1                   Michael      5
## 5     Person 1               Toy Story 5      4
## 6     Person 1                 Obsession      4
## 7     Person 2 Spider-Man: Brand New Day      4
## 8     Person 2                   Sinners      5
## 9     Person 2               The Odyssey      4
## 10    Person 2                   Michael      5
## 11    Person 2               Toy Story 5     NA
## 12    Person 2                 Obsession      3
## 13    Person 3 Spider-Man: Brand New Day      5
## 14    Person 3                   Sinners     NA
## 15    Person 3               The Odyssey      5
## 16    Person 3                   Michael      4
## 17    Person 3               Toy Story 5      2
## 18    Person 3                 Obsession      5
## 19    Person 4 Spider-Man: Brand New Day      4
## 20    Person 4                   Sinners      4
## 21    Person 4               The Odyssey     NA
## 22    Person 4                   Michael      4
## 23    Person 4               Toy Story 5      3
## 24    Person 4                 Obsession     NA
## 25    Person 5 Spider-Man: Brand New Day      5
## 26    Person 5                   Sinners      5
## 27    Person 5               The Odyssey      3
## 28    Person 5                   Michael      5
## 29    Person 5               Toy Story 5     NA
## 30    Person 5                 Obsession     NA

Analysis

First, I calculated the global mean, which represents the average of all available movie ratings. Missing ratings were excluded from this calculation.

global_mean <- mean(ratings_df$rating, na.rm = TRUE)

global_mean
## [1] 4.217391

Next, I calculated the average rating for each participant. The user effect represents the difference between each participant’s average rating and the global mean.

user_effects <- ratings_df %>%
  group_by(participant) %>%
  summarise(
    user_average = mean(rating, na.rm = TRUE)
  ) %>%
  mutate(
    user_effect = user_average - global_mean
  )

user_effects
## # A tibble: 5 × 3
##   participant user_average user_effect
##   <chr>              <dbl>       <dbl>
## 1 Person 1            4.4       0.183 
## 2 Person 2            4.2      -0.0174
## 3 Person 3            4.2      -0.0174
## 4 Person 4            3.75     -0.467 
## 5 Person 5            4.5       0.283

I then calculated the average rating for each movie. The movie effect represents the difference between each movie’s average rating and the global mean.

movie_effects <- ratings_df %>%
  group_by(movie) %>%
  summarise(
    movie_average = mean(rating, na.rm = TRUE)
  ) %>%
  mutate(
    movie_effect = movie_average - global_mean
  )

movie_effects
## # A tibble: 6 × 3
##   movie                     movie_average movie_effect
##   <chr>                             <dbl>        <dbl>
## 1 Michael                             4.6        0.383
## 2 Obsession                           4         -0.217
## 3 Sinners                             4.5        0.283
## 4 Spider-Man: Brand New Day           4.6        0.383
## 5 The Odyssey                         4         -0.217
## 6 Toy Story 5                         3         -1.22

Since the goal is to predict ratings for movies that participants have not seen, I identified the observations with missing ratings.

unrated_movies <- ratings_df %>%
  filter(is.na(rating))

unrated_movies
##   participant       movie rating
## 1    Person 1 The Odyssey     NA
## 2    Person 2 Toy Story 5     NA
## 3    Person 3     Sinners     NA
## 4    Person 4 The Odyssey     NA
## 5    Person 4   Obsession     NA
## 6    Person 5 Toy Story 5     NA
## 7    Person 5   Obsession     NA

I calculated the Global Baseline Estimate for each missing rating by adding the global mean, the participant’s user effect, and the movie effect.

baseline_predictions <- unrated_movies %>%
  left_join(user_effects, by = "participant") %>%
  left_join(movie_effects, by = "movie") %>%
  mutate(
    predicted_rating = global_mean + user_effect + movie_effect
  )

baseline_predictions
##   participant       movie rating user_average user_effect movie_average
## 1    Person 1 The Odyssey     NA         4.40   0.1826087           4.0
## 2    Person 2 Toy Story 5     NA         4.20  -0.0173913           3.0
## 3    Person 3     Sinners     NA         4.20  -0.0173913           4.5
## 4    Person 4 The Odyssey     NA         3.75  -0.4673913           4.0
## 5    Person 4   Obsession     NA         3.75  -0.4673913           4.0
## 6    Person 5 Toy Story 5     NA         4.50   0.2826087           3.0
## 7    Person 5   Obsession     NA         4.50   0.2826087           4.0
##   movie_effect predicted_rating
## 1   -0.2173913         4.182609
## 2   -1.2173913         2.982609
## 3    0.2826087         4.482609
## 4   -0.2173913         3.532609
## 5   -0.2173913         3.532609
## 6   -1.2173913         3.282609
## 7   -0.2173913         4.282609

I selected the movie with the highest predicted rating for each participant as the recommendation.

recommendations <- baseline_predictions %>%
  group_by(participant) %>%
  slice_max(
    order_by = predicted_rating,
    n = 1, 
    with_ties = FALSE
  ) %>%
  select(participant, movie, predicted_rating)

recommendations
## # A tibble: 5 × 3
## # Groups:   participant [5]
##   participant movie       predicted_rating
##   <chr>       <chr>                  <dbl>
## 1 Person 1    The Odyssey             4.18
## 2 Person 2    Toy Story 5             2.98
## 3 Person 3    Sinners                 4.48
## 4 Person 4    The Odyssey             3.53
## 5 Person 5    Obsession               4.28

The Global Baseline Estimate produced one movie recommendation for each participant based on their predicted ratings. Person 1 was recommended The Odyssey with a predicted rating of 4.18, Person 2 was recommended Toy Story 5 with a predicted rating of 2.98, Person 3 was recommended Sinners with a predicted rating of 4.48, Person 4 was recommended The Odyssey with a predicted rating of 3.53, and Person 5 was recommended Obsession with a predicted rating of 4.28.

Conclusion

The Global Baseline Estimate allowed me to use the existing movie ratings to predict how participants might rate movies they had not seen. The method uses the overall average rating along with user and movie effects to account for differences in how participants rate movies and how movies are rated overall. The predicted ratings were used to generate movie recommendations for each participant.