For this project, I will use the movie ratings data collected in the previous assignment (LAB2A) to build a simple recommendation system using the Global Baseline Estimate (GBE) algorithm. The dataset consists of ratings from multiple users who rated several recent movies on a scale from 1 to 5. Because not every participant rated every movie, the dataset contains missing values that represent movies a user has not seen. To tackle this problem, I will first import the movie ratings data into R and organize it into a user-movie ratings matrix. Next, I will calculate the global average rating across all movies and users. After determining the global average, I will calculate user-specific biases and movie-specific biases. User bias measures whether a user generally rates movies higher or lower than average, while movie bias measures whether a movie tends to receive higher or lower ratings than average.Using these values, I will implement the Global Baseline Estimate algorithm to predict ratings for movies that users have not yet seen. The predicted ratings will then be used to generate recommendations. For each user, the movie with the highest predicted rating among the unseen movies will be identified as the recommended movie. Finally, I will evaluate the results and discuss whether the recommendations appear reasonable based on the users’ existing rating patterns. I will also summarize the strengths and limitations of using a non-personalized recommendation algorithm such as Global Baseline Estimates. Anticipated Challenges
in this session, I imported a the data set from the previous assignment 2A. The dataset contains user ratings for several recently released movies on a scale from 1 to 5. Because not every user rated every movie, some ratings are missing.To prepare the data for the Global Baseline Estimate algorithm, duplicate records were removed, and the dataset was transformed from a long format into a user-movie ratings matrix. In this format, each row represents a user, and each column represents a movie. Missing values (NA) indicate movies that a user has not yet seen.
library(readr)
movie_ratings <- read_csv("https://raw.githubusercontent.com/lioneljr17/LDATA607/refs/heads/main/607LAB2A_files/movie_ratings.csv")
## Rows: 56 Columns: 4
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (3): users_name, movie_title, rating_status
## dbl (1): rating
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
movies_ratings_withoutdup<- movie_ratings%>%
distinct(users_name,movie_title,.keep_all = TRUE)
#View(movie_ratings)
head((movie_ratings))
## # A tibble: 6 × 4
## users_name movie_title rating rating_status
## <chr> <chr> <dbl> <chr>
## 1 aly Michael 5 5
## 2 aly Michael 5 5
## 3 aly Minions & Monsters 5 5
## 4 aly spider man brand new day 5 5
## 5 aly The Odyssey NA Not Seen
## 6 aly The Odyssey NA Not Seen
Movie_Rating_view <- movie_ratings %>%
select (users_name, movie_title, rating)%>%
distinct() %>%
pivot_wider (names_from =movie_title, values_from = rating)
names(Movie_Rating_view) <- gsub(" ","_", names(Movie_Rating_view))
names(Movie_Rating_view) <- gsub("&","and", names(Movie_Rating_view))
head(Movie_Rating_view)
## # A tibble: 6 × 7
## users_name Michael Minions_and_Monsters spider_man_brand_new_day The_Odyssey
## <chr> <dbl> <dbl> <dbl> <dbl>
## 1 aly 5 5 5 NA
## 2 church 4 NA 5 NA
## 3 lionel NA 3 5 4
## 4 MIKE 3 NA 4 5
## 5 paul NA 5 5 NA
## 6 zege 2 NA 5 NA
## # ℹ 2 more variables: The_Super_Mario_Galaxy_Movie <dbl>, Toy_Story_5 <dbl>
in this session I procceed with the step which is to calculate the global average rating. This value represents the average rating across all users and movies in the dataset.The global average serves as the baseline prediction for the recommendation system. All future predictions begin with this value and are adjusted using user and movie biases.
global_avrg<- mean (movies_ratings_withoutdup$rating,na.rm = TRUE)
global_avrg
## [1] 4.15
After calculating the global average, the next step is to determine how each user typically rates movies. A user average is calculated by taking the mean of all ratings provided by a specific user. User bias measures whether a user tends to rate movies higher or lower than the overall average. A positive user bias indicates that a user generally gives higher ratings, while a negative user bias indicates that a user tends to rate movies lower than average.
Movie_Rating_view$users_avrg <- rowMeans(Movie_Rating_view[-1], na.rm = TRUE)
Movie_Rating_view
## # A tibble: 6 × 8
## users_name Michael Minions_and_Monsters spider_man_brand_new_day The_Odyssey
## <chr> <dbl> <dbl> <dbl> <dbl>
## 1 aly 5 5 5 NA
## 2 church 4 NA 5 NA
## 3 lionel NA 3 5 4
## 4 MIKE 3 NA 4 5
## 5 paul NA 5 5 NA
## 6 zege 2 NA 5 NA
## # ℹ 3 more variables: The_Super_Mario_Galaxy_Movie <dbl>, Toy_Story_5 <dbl>,
## # users_avrg <dbl>
Movie_Rating_view$users_bias <- Movie_Rating_view$users_avrg - global_avrg
Movie_Rating_view
## # A tibble: 6 × 9
## users_name Michael Minions_and_Monsters spider_man_brand_new_day The_Odyssey
## <chr> <dbl> <dbl> <dbl> <dbl>
## 1 aly 5 5 5 NA
## 2 church 4 NA 5 NA
## 3 lionel NA 3 5 4
## 4 MIKE 3 NA 4 5
## 5 paul NA 5 5 NA
## 6 zege 2 NA 5 NA
## # ℹ 4 more variables: The_Super_Mario_Galaxy_Movie <dbl>, Toy_Story_5 <dbl>,
## # users_avrg <dbl>, users_bias <dbl>
The recommendation system must also account for differences in movie popularity. Some movies tend to receive higher ratings from most users, while others receive lower ratings.To capture this effect, a movie average is calculated for each movie. Movie bias is then computed by subtracting the global average from the movie average. A positive movie bias indicates that a movie is generally well-liked, while a negative movie bias suggests that the movie receives lower ratings than average.
movie_avrg <- colMeans(Movie_Rating_view[,2:(ncol(Movie_Rating_view))], na.rm = TRUE)
movie_avrg
## Michael Minions_and_Monsters
## 3.500000 4.333333
## spider_man_brand_new_day The_Odyssey
## 4.833333 4.500000
## The_Super_Mario_Galaxy_Movie Toy_Story_5
## 4.000000 3.500000
## users_avrg users_bias
## 4.125000 -0.025000
movie_avrg_row<- data.frame(users_name ="movie_avrg", t(movie_avrg))
movie_avrg_row
## users_name Michael Minions_and_Monsters spider_man_brand_new_day The_Odyssey
## 1 movie_avrg 3.5 4.333333 4.833333 4.5
## The_Super_Mario_Galaxy_Movie Toy_Story_5 users_avrg users_bias
## 1 4 3.5 4.125 -0.025
Movie_Rating_view <- bind_rows(Movie_Rating_view,movie_avrg_row)
movie_bias <- movie_avrg - global_avrg
movie_bias
## Michael Minions_and_Monsters
## -0.6500000 0.1833333
## spider_man_brand_new_day The_Odyssey
## 0.6833333 0.3500000
## The_Super_Mario_Galaxy_Movie Toy_Story_5
## -0.1500000 -0.6500000
## users_avrg users_bias
## -0.0250000 -4.1750000
movie_bias_row<- data.frame(users_name ="movie_bias", t(movie_bias))
movie_bias_row
## users_name Michael Minions_and_Monsters spider_man_brand_new_day The_Odyssey
## 1 movie_bias -0.65 0.1833333 0.6833333 0.35
## The_Super_Mario_Galaxy_Movie Toy_Story_5 users_avrg users_bias
## 1 -0.15 -0.65 -0.025 -4.175
Movie_Rating_view <- bind_rows(Movie_Rating_view,movie_bias_row)
The final step of the Global Baseline Estimate algorithm is to predict ratings for movies that users have not yet seen. Missing ratings are represented by NA values in the dataset.The algorithm scans each user and identifies movies that have not been rated. A predicted rating is then generated using the formula above and inserted into the recommendation table.
Prediction_table <- Movie_Rating_view %>%
filter(!users_name %in% c("movie_avrg","movie_bias"))
#movie_cols <-2.7
movie_cols <- names(Movie_Rating_view)[!(names(Movie_Rating_view)%in% c("users_name","users_bias","users_avrg"))]
movie_cols
## [1] "Michael" "Minions_and_Monsters"
## [3] "spider_man_brand_new_day" "The_Odyssey"
## [5] "The_Super_Mario_Galaxy_Movie" "Toy_Story_5"
global_avrg + Movie_Rating_view$users_bias +movie_bias["The_Odyssey"]
## [1] 4.850000 5.016667 4.600000 4.350000 5.016667 3.016667 4.475000 0.325000
for(i in 1:nrow(Prediction_table)) {
current_users_bias <- Prediction_table$users_bias[i]
for(j in movie_cols) {
if(!is.na(Prediction_table[[j]][i])) {
Prediction_table[i, j] <- NA
}
else{
cat("\n----------------------\n")
cat("User:", Prediction_table$users_name[i], "\n")
cat("Movie:", j, "\n")
print(global_avrg)
print(current_users_bias)
print(movie_bias[j])
predicted_rating <-
global_avrg +
current_users_bias +
movie_bias[j]
cat("Prediction =", predicted_rating, "\n")
Prediction_table[[j]][i] <-
round(predicted_rating, 2)
}
}
}
##
## ----------------------
## User: aly
## Movie: The_Odyssey
## [1] 4.15
## [1] 0.35
## The_Odyssey
## 0.35
## Prediction = 4.85
##
## ----------------------
## User: aly
## Movie: The_Super_Mario_Galaxy_Movie
## [1] 4.15
## [1] 0.35
## The_Super_Mario_Galaxy_Movie
## -0.15
## Prediction = 4.35
##
## ----------------------
## User: church
## Movie: Minions_and_Monsters
## [1] 4.15
## [1] 0.5166667
## Minions_and_Monsters
## 0.1833333
## Prediction = 4.85
##
## ----------------------
## User: church
## Movie: The_Odyssey
## [1] 4.15
## [1] 0.5166667
## The_Odyssey
## 0.35
## Prediction = 5.016667
##
## ----------------------
## User: church
## Movie: The_Super_Mario_Galaxy_Movie
## [1] 4.15
## [1] 0.5166667
## The_Super_Mario_Galaxy_Movie
## -0.15
## Prediction = 4.516667
##
## ----------------------
## User: lionel
## Movie: Michael
## [1] 4.15
## [1] 0.1
## Michael
## -0.65
## Prediction = 3.6
##
## ----------------------
## User: lionel
## Movie: The_Super_Mario_Galaxy_Movie
## [1] 4.15
## [1] 0.1
## The_Super_Mario_Galaxy_Movie
## -0.15
## Prediction = 4.1
##
## ----------------------
## User: MIKE
## Movie: Minions_and_Monsters
## [1] 4.15
## [1] -0.15
## Minions_and_Monsters
## 0.1833333
## Prediction = 4.183333
##
## ----------------------
## User: MIKE
## Movie: The_Super_Mario_Galaxy_Movie
## [1] 4.15
## [1] -0.15
## The_Super_Mario_Galaxy_Movie
## -0.15
## Prediction = 3.85
##
## ----------------------
## User: MIKE
## Movie: Toy_Story_5
## [1] 4.15
## [1] -0.15
## Toy_Story_5
## -0.65
## Prediction = 3.35
##
## ----------------------
## User: paul
## Movie: Michael
## [1] 4.15
## [1] 0.5166667
## Michael
## -0.65
## Prediction = 4.016667
##
## ----------------------
## User: paul
## Movie: The_Odyssey
## [1] 4.15
## [1] 0.5166667
## The_Odyssey
## 0.35
## Prediction = 5.016667
##
## ----------------------
## User: paul
## Movie: Toy_Story_5
## [1] 4.15
## [1] 0.5166667
## Toy_Story_5
## -0.65
## Prediction = 4.016667
##
## ----------------------
## User: zege
## Movie: Minions_and_Monsters
## [1] 4.15
## [1] -1.483333
## Minions_and_Monsters
## 0.1833333
## Prediction = 2.85
##
## ----------------------
## User: zege
## Movie: The_Odyssey
## [1] 4.15
## [1] -1.483333
## The_Odyssey
## 0.35
## Prediction = 3.016667
##
## ----------------------
## User: zege
## Movie: The_Super_Mario_Galaxy_Movie
## [1] 4.15
## [1] -1.483333
## The_Super_Mario_Galaxy_Movie
## -0.15
## Prediction = 2.516667
The Global Baseline Estimate algorithm successfully generated predicted ratings for movies that users had not previously rated and seen. These predictions were based on the overall average rating, user rating tendencies, and movie popularity trends. Users with positive biases generally received higher predicted ratings, while users with negative biases received lower predicted ratings. Similarly, movies with positive movie biases tended to receive higher recommendation scores.The completed prediction table provides estimated ratings for all previously missing user-movie combinations and serves as the foundation for generating movie recommendations.
In conclusion, The recommendation process began by calculating a global average rating, followed by user and movie biases. These values were then used to estimate ratings for movies that users had not yet seen. The resulting predictions provide a simple but effective recommendation strategy without requiring advanced collaborative filtering techniques. One limitation of this approach is that the dataset contains a relatively small number of users and movies. As a result, recommendations may be heavily influenced by individual ratings. However, the project successfully demonstrates the core concepts behind recommendation systems and provides a foundation for more advanced recommendation algorithms in future assignments.