Global Baseline Estimates

Global Baseline Estimates

Introduction

The simplest way to recommend a movie is based on the global baseline estimate. We will take the mean movie rating for all movies add the relative rating of the movie we are looking to evaluate for recommendation, and finally add the average rating of the particular user we are looking to recommend for. The current average of all movies being evaluated in the dataset is 3.93 and we can use this as a universal constant.

The Formula

The formula being used is in the below code snippet. We will take the universal average rating names ua for universal average, then take the user average score in the 18th column, and finally take the users personal movie average minus the mean movie average.

In order to find which movies to be evaluated, we will get the index of all NA values in the original dataset for each user so we know which movies need a recommendation score. Then we will cap the score at 5 so it can be more akin to a regular user score.

for( i in 2:(length(df)-2)){
  
  naValues<-which(is.na(df[,i]))
  for( j in 1:length(naValues)){
    if ( (ua + df[18,i] + df[naValues[j],9] ) <= 5 ){
      estimatedf[naValues[j],i] <- (ua + df[18,i] + df[naValues[j],9] )
    } else { 
      estimatedf[naValues[j],i] <- 5
    }
  }
  
}

Recommendation score table

We are left with a table with all of the recommendation scores in the same style as the original dataset

Global Baseline scores
Critic CaptainAmerica Deadpool Frozen JungleBook PitchPerfect2 StarWarsForce user.avg user.avg...mean.movie
Burton 4.34 4.51 3.79 0.00 2.78 0.00 4.00 0.07
Charley 0.00 0.00 0.00 0.00 0.00 0.00 3.50 -0.43
Dan 5.00 0.00 4.79 4.97 3.78 0.00 5.00 1.07
Dieudonne 0.00 0.00 4.45 4.63 3.44 0.00 4.67 0.73
Matt 0.00 3.76 0.00 3.22 0.00 0.00 3.25 -0.68
Mauricio 0.00 4.01 0.00 0.00 0.00 3.72 3.50 -0.43
Max 0.00 0.00 0.00 0.00 0.00 0.00 3.33 -0.60
Nathan 4.34 4.51 3.79 3.97 2.78 0.00 4.00 0.07
Param 0.00 0.00 0.00 3.47 2.28 0.00 3.50 -0.43
Parshu 0.00 0.00 0.00 0.00 0.00 0.00 3.67 -0.27
Prashanth 0.00 0.00 0.00 0.00 3.58 0.00 4.80 0.87
Shipra 4.34 4.51 0.00 0.00 2.78 0.00 4.00 0.07
Sreejaya 0.00 0.00 0.00 0.00 0.00 0.00 4.67 0.73
Steve 0.00 4.51 3.79 3.97 2.78 0.00 4.00 0.07
Vuthy 0.00 0.00 0.00 0.00 0.00 3.82 3.60 -0.33
Xingjia 5.00 5.00 0.00 0.00 3.78 5.00 5.00 1.07
movie avg 4.27 4.44 3.73 3.90 2.71 4.15 3.93 NA
movie avg - mean movie 0.34 0.51 -0.21 -0.03 -1.22 0.22 NA NA

Recommendations

At this point we can create a dataframe that has the names of each critic and extract the name of the movie recommended as well as the predicted score and display it as a recommendation. All critics that have seen all provided movies will be removed.

names <- df$Critic[1:16]
movies <-  vector( mode = "character", 16)
pred.score <- vector( mode = "numeric", 16)

recdf <- data.frame(Critic = names, 
                    movie.rec = movies, 
                    predicted_score = pred.score)

for( i in 1:length(recdf$Critic)){
  
  recdf[i,3] <- max(estimatedf[i,2:6])
  recdf[i,2] <- names(which.max(estimatedf[i,2:6]))
}

recdf <- subset(recdf, predicted_score > 0 )
movie recs
Critic movie.rec predicted_score
Burton Deadpool 4.51
Dan CaptainAmerica 5.00
Dieudonne JungleBook 4.63
Matt Deadpool 3.76
Mauricio Deadpool 4.01
Nathan Deadpool 4.51
Param JungleBook 3.47
Prashanth PitchPerfect2 3.58
Shipra Deadpool 4.51
Steve Deadpool 4.51
Xingjia CaptainAmerica 5.00