The Data and Packages

The data for this plot come from the training data for a Kaggle competition to predict movie revenue from about two years ago. The data can be downloaded here:

Data Link

Load packages:

suppressMessages(library(tidyverse))
suppressMessages(library(plotly))
suppressMessages(library(lubridate))

Preparing the data

  • The original data set contains information on 23 variables about 30000 movies.
  • Subset the data to focus on high-revenue movies (at least $10 million)
  • Convert dates to date type and extract the year
Movies <- as_tibble(read.csv(
     "~/Desktop/tmdb-box-office-prediction/train.csv"))
Movies <- Movies %>% 
     select(title,budget,popularity,revenue,release_date) %>% 
     filter((revenue >= 1e7) & (budget >0))
Movies <- Movies %>% 
     mutate(Year=year(as.POSIXlt(release_date,format="%m/%d/%Y"))) %>%
     mutate(Year=Year+1900*(Year >= 21)+2000*(Year < 21))

Create the Plot

  • \(x\)-axis is budget, \(y\)-axis is revenue
  • Markers are labeled with movie title on hover with mouse pointer
  • Color is determine by release year
  • Size of data markers reflects “popularity” (I’m not sure how this was measured, but it’s interesting)
  • The warning produced is apparently a known bug when using markers with variable size

The Plot

## Warning: `line.width` does not currently support multiple values.

Made on 9 Feb 2021