In this assignment, I will use the movie ratings data I collected from five participants for the six selected movies. The dataset contains ratings on a scale from 1 to 5, along with missing values for movies that participants had not seen. I will use this data to create a recommendation system using the Global Baseline Estimate method. I will use my ratings data to calculate user and movie effects and predict ratings for movies that participants have not seen. The movie with the highest predicted rating will be used as the recommendation for each participant.
One challenge I anticipate is calculating the Global Baseline Estimates correctly when participants have rated different numbers of movies. I will use the available ratings to calculate the overall, participant, and movie averages before using the model to predict the missing ratings.
Data Source: Primary movie rating data collected from five participants from the previous Movie Ratings assignment.
library(DBI)
## Warning: package 'DBI' was built under R version 4.4.3
library(RPostgres)
## Warning: package 'RPostgres' was built under R version 4.4.3
library(dplyr)
## Warning: package 'dplyr' was built under R version 4.4.3
##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
con <- dbConnect(
RPostgres::Postgres(),
dbname = "movie_ratings",
host = "localhost",
port = 5432,
user = "postgres",
password = rstudioapi::askForPassword("Enter PostgreSQL password")
)
# I checked the tables stored in the database.
dbListTables(con)
## [1] "movies" "participants" "ratings"
I used a SQL query to join the three tables and load the results into an R dataframe.
ratings_df <- dbGetQuery(
con,
"
SELECT
p.name AS participant,
m.title AS movie,
r.rating
FROM ratings AS r
JOIN participants AS p
ON r.user_id = p.user_id
JOIN movies AS m
ON r.movie_id = m.movie_id
ORDER BY p.user_id, m.movie_id;
"
)
ratings_df
## participant movie rating
## 1 Person 1 Spider-Man: Brand New Day 5
## 2 Person 1 Sinners 4
## 3 Person 1 The Odyssey NA
## 4 Person 1 Michael 5
## 5 Person 1 Toy Story 5 4
## 6 Person 1 Obsession 4
## 7 Person 2 Spider-Man: Brand New Day 4
## 8 Person 2 Sinners 5
## 9 Person 2 The Odyssey 4
## 10 Person 2 Michael 5
## 11 Person 2 Toy Story 5 NA
## 12 Person 2 Obsession 3
## 13 Person 3 Spider-Man: Brand New Day 5
## 14 Person 3 Sinners NA
## 15 Person 3 The Odyssey 5
## 16 Person 3 Michael 4
## 17 Person 3 Toy Story 5 2
## 18 Person 3 Obsession 5
## 19 Person 4 Spider-Man: Brand New Day 4
## 20 Person 4 Sinners 4
## 21 Person 4 The Odyssey NA
## 22 Person 4 Michael 4
## 23 Person 4 Toy Story 5 3
## 24 Person 4 Obsession NA
## 25 Person 5 Spider-Man: Brand New Day 5
## 26 Person 5 Sinners 5
## 27 Person 5 The Odyssey 3
## 28 Person 5 Michael 5
## 29 Person 5 Toy Story 5 NA
## 30 Person 5 Obsession NA
First, I calculated the global mean, which represents the average of all available movie ratings. Missing ratings were excluded from this calculation.
global_mean <- mean(ratings_df$rating, na.rm = TRUE)
global_mean
## [1] 4.217391
Next, I calculated the average rating for each participant. The user effect represents the difference between each participant’s average rating and the global mean.
user_effects <- ratings_df %>%
group_by(participant) %>%
summarise(
user_average = mean(rating, na.rm = TRUE)
) %>%
mutate(
user_effect = user_average - global_mean
)
user_effects
## # A tibble: 5 × 3
## participant user_average user_effect
## <chr> <dbl> <dbl>
## 1 Person 1 4.4 0.183
## 2 Person 2 4.2 -0.0174
## 3 Person 3 4.2 -0.0174
## 4 Person 4 3.75 -0.467
## 5 Person 5 4.5 0.283
I then calculated the average rating for each movie. The movie effect represents the difference between each movie’s average rating and the global mean.
movie_effects <- ratings_df %>%
group_by(movie) %>%
summarise(
movie_average = mean(rating, na.rm = TRUE)
) %>%
mutate(
movie_effect = movie_average - global_mean
)
movie_effects
## # A tibble: 6 × 3
## movie movie_average movie_effect
## <chr> <dbl> <dbl>
## 1 Michael 4.6 0.383
## 2 Obsession 4 -0.217
## 3 Sinners 4.5 0.283
## 4 Spider-Man: Brand New Day 4.6 0.383
## 5 The Odyssey 4 -0.217
## 6 Toy Story 5 3 -1.22
Since the goal is to predict ratings for movies that participants have not seen, I identified the observations with missing ratings.
unrated_movies <- ratings_df %>%
filter(is.na(rating))
unrated_movies
## participant movie rating
## 1 Person 1 The Odyssey NA
## 2 Person 2 Toy Story 5 NA
## 3 Person 3 Sinners NA
## 4 Person 4 The Odyssey NA
## 5 Person 4 Obsession NA
## 6 Person 5 Toy Story 5 NA
## 7 Person 5 Obsession NA
I calculated the Global Baseline Estimate for each missing rating by adding the global mean, the participant’s user effect, and the movie effect.
baseline_predictions <- unrated_movies %>%
left_join(user_effects, by = "participant") %>%
left_join(movie_effects, by = "movie") %>%
mutate(
predicted_rating = global_mean + user_effect + movie_effect
)
baseline_predictions
## participant movie rating user_average user_effect movie_average
## 1 Person 1 The Odyssey NA 4.40 0.1826087 4.0
## 2 Person 2 Toy Story 5 NA 4.20 -0.0173913 3.0
## 3 Person 3 Sinners NA 4.20 -0.0173913 4.5
## 4 Person 4 The Odyssey NA 3.75 -0.4673913 4.0
## 5 Person 4 Obsession NA 3.75 -0.4673913 4.0
## 6 Person 5 Toy Story 5 NA 4.50 0.2826087 3.0
## 7 Person 5 Obsession NA 4.50 0.2826087 4.0
## movie_effect predicted_rating
## 1 -0.2173913 4.182609
## 2 -1.2173913 2.982609
## 3 0.2826087 4.482609
## 4 -0.2173913 3.532609
## 5 -0.2173913 3.532609
## 6 -1.2173913 3.282609
## 7 -0.2173913 4.282609
I selected the movie with the highest predicted rating for each participant as the recommendation.
recommendations <- baseline_predictions %>%
group_by(participant) %>%
slice_max(
order_by = predicted_rating,
n = 1,
with_ties = FALSE
) %>%
select(participant, movie, predicted_rating)
recommendations
## # A tibble: 5 × 3
## # Groups: participant [5]
## participant movie predicted_rating
## <chr> <chr> <dbl>
## 1 Person 1 The Odyssey 4.18
## 2 Person 2 Toy Story 5 2.98
## 3 Person 3 Sinners 4.48
## 4 Person 4 The Odyssey 3.53
## 5 Person 5 Obsession 4.28
The Global Baseline Estimate produced one movie recommendation for each participant based on their predicted ratings. Person 1 was recommended The Odyssey with a predicted rating of 4.18, Person 2 was recommended Toy Story 5 with a predicted rating of 2.98, Person 3 was recommended Sinners with a predicted rating of 4.48, Person 4 was recommended The Odyssey with a predicted rating of 3.53, and Person 5 was recommended Obsession with a predicted rating of 4.28.
The Global Baseline Estimate allowed me to use the existing movie ratings to predict how participants might rate movies they had not seen. The method uses the overall average rating along with user and movie effects to account for differences in how participants rate movies and how movies are rated overall. The predicted ratings were used to generate movie recommendations for each participant.