Romance analysis

Imports

Scores extraction

As for now there are 4232 annotated rows.

# General love score
movies_progress$gen_love_score <- as.numeric(str_extract(movies_progress$love_gpt, "(?<=\\/ SCORE = )\\d+|NA"))
Warning: NAs introduits lors de la conversion automatique

Validity checks

Comparing average of General Love Score of romance movies vs non romance movies:

# Compare the average of both scores between works categorized under the romance IMDb genre and non-Romance genres. The expectation is that this difference will be significant with a t-test.
# General Love Score
# Filter
gen_love_score_romance <- movies_progress$gen_love_score[movies_progress$romance == 1]
gen_love_score_non_romance <- movies_progress$gen_love_score[movies_progress$romance == 0]

# t test
t.test(gen_love_score_romance, gen_love_score_non_romance, mu=0)

    Welch Two Sample t-test

data:  gen_love_score_romance and gen_love_score_non_romance
t = 50.044, df = 1173.4, p-value < 2.2e-16
alternative hypothesis: true difference in means is not equal to 0
95 percent confidence interval:
 4.164561 4.504430
sample estimates:
mean of x mean of y 
 7.257019  2.922524 

For the General Love Score, their is a significant scores difference between romance and non romance movies (p < 2.2e-16).

Statistical analysis

General Love Score

Models

# A function to include AIC in the upcoming model summaries
summary_AIC <- function(model) {
  summary_output <- summary(model)
  aic <- AIC(model)
  summary_output$aic <- aic
  print(summary_output)
}
1

One linear model with the year as the explanatory variable:

lm1 <- lm(gen_love_score ~ year, data=movies_progress)
summary_AIC(lm1)$aic

Call:
lm(formula = gen_love_score ~ year, data = movies_progress)

Residuals:
    Min      1Q  Median      3Q     Max 
-3.6909 -1.6730 -0.6612  1.4517  6.5408 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)  
(Intercept) -8.301847   5.153771  -1.611   0.1073  
year         0.005940   0.002576   2.306   0.0211 *
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 2.66 on 6074 degrees of freedom
  (4923 observations effacées parce que manquantes)
Multiple R-squared:  0.0008748, Adjusted R-squared:  0.0007103 
F-statistic: 5.318 on 1 and 6074 DF,  p-value: 0.02114
[1] 29134.74
2

One linear model with the year as the explanatory variable and the number of movies published that year as the control:

# Number of movies published per year
movies_progress_noNA <- movies_progress %>%
  filter(gen_love_score != "NA") %>%
  group_by(year) %>%
  mutate(movies_per_year = n()) %>%
  ungroup()

lm2 <- lm(gen_love_score ~ year + movies_per_year, data=movies_progress_noNA)
summary_AIC(lm2)$aic

Call:
lm(formula = gen_love_score ~ year + movies_per_year, data = movies_progress_noNA)

Residuals:
    Min      1Q  Median      3Q     Max 
-3.7392 -1.6641 -0.6581  1.4464  6.7667 

Coefficients:
                  Estimate Std. Error t value Pr(>|t|)  
(Intercept)     11.6122633 12.9099479   0.899   0.3684  
year            -0.0041934  0.0065507  -0.640   0.5221  
movies_per_year  0.0013130  0.0007805   1.682   0.0925 .
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 2.659 on 6073 degrees of freedom
Multiple R-squared:  0.00134,   Adjusted R-squared:  0.001011 
F-statistic: 4.075 on 2 and 6073 DF,  p-value: 0.01704
[1] 29133.91
3

One quadratic model with year squared as the explanatory variable:

lm3 <- lm(gen_love_score ~ year + year^2, data=movies_progress)
summary_AIC(lm3)$aic

Call:
lm(formula = gen_love_score ~ year + year^2, data = movies_progress)

Residuals:
    Min      1Q  Median      3Q     Max 
-3.6909 -1.6730 -0.6612  1.4517  6.5408 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)  
(Intercept) -8.301847   5.153771  -1.611   0.1073  
year         0.005940   0.002576   2.306   0.0211 *
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 2.66 on 6074 degrees of freedom
  (4923 observations effacées parce que manquantes)
Multiple R-squared:  0.0008748, Adjusted R-squared:  0.0007103 
F-statistic: 5.318 on 1 and 6074 DF,  p-value: 0.02114
[1] 29134.74
4

One quadratic model with year squared as the explanatory variable and the number of movies published that year as the control:

lm4 <- lm(gen_love_score ~ year + year^2 + movies_per_year, data=movies_progress_noNA)
summary_AIC(lm4)$aic

Call:
lm(formula = gen_love_score ~ year + year^2 + movies_per_year, 
    data = movies_progress_noNA)

Residuals:
    Min      1Q  Median      3Q     Max 
-3.7392 -1.6641 -0.6581  1.4464  6.7667 

Coefficients:
                  Estimate Std. Error t value Pr(>|t|)  
(Intercept)     11.6122633 12.9099479   0.899   0.3684  
year            -0.0041934  0.0065507  -0.640   0.5221  
movies_per_year  0.0013130  0.0007805   1.682   0.0925 .
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 2.659 on 6073 degrees of freedom
Multiple R-squared:  0.00134,   Adjusted R-squared:  0.001011 
F-statistic: 4.075 on 2 and 6073 DF,  p-value: 0.01704
[1] 29133.91

Models comparison

Comparing regression output:

stargazer(lm1, lm2, lm3, lm4, type="html", digits=2, title="Regression models comparison for General Love Score", style="qje", column.labels=c("OLS", "OLS quadratic", "OLS + control", "OLS + control + quadratic"), dep.var.labels = "General Love score")
Regression models comparison for General Love Score
General Love score
OLS OLS quadratic OLS + control OLS + control + quadratic
(1) (2) (3) (4)
year 0.01** -0.004 0.01** -0.004
(0.003) (0.01) (0.003) (0.01)
movies_per_year 0.001* 0.001*
(0.001) (0.001)
Constant -8.30 11.61 -8.30 11.61
(5.15) (12.91) (5.15) (12.91)
N 6,076 6,076 6,076 6,076
R2 0.001 0.001 0.001 0.001
Adjusted R2 0.001 0.001 0.001 0.001
Residual Std. Error 2.66 (df = 6074) 2.66 (df = 6073) 2.66 (df = 6074) 2.66 (df = 6073)
F Statistic 5.32** (df = 1; 6074) 4.08** (df = 2; 6073) 5.32** (df = 1; 6074) 4.08** (df = 2; 6073)
Notes: ***Significant at the 1 percent level.
**Significant at the 5 percent level.
*Significant at the 10 percent level.

Comparing AIC scores:

[1] "Model 1 AIC: 29134.7409617425"
[1] "Model 2 AIC: 29133.9098603473"
[1] "Model 3 AIC: 29134.7409617425"
[1] "Model 4 AIC: 29133.9098603473"

Romantic Love Score

One linear model with the year as the explanatory variable:

One linear model with the year as the explanatory variable and the number of movies published that year as the control:

One quadratic model with year squared as the explanatory variable:

One quadratic model with year squared as the explanatory variable and the number of movies published that year as the control:

Plots

General Love scores

movies_progress_noNA <- movies_progress_noNA %>%
  group_by(year) %>%
  mutate(mean_gen_love_score = (mean(gen_love_score)))
  
gen_love_plot <- movies_progress_noNA %>%
  ggplot(aes(x=as.numeric(year), y=mean_gen_love_score)) + geom_point() +
  labs(
    title="Evolution of General Love scores",
    xlab="Year",
    ylab="Mean score"
  ) +
  theme_bw()
  
gen_love_plot

Romantic love scores