This code through explores how R can be used to analyze movie genres and worldwide box office revenue. Using a real movie dataset, I will demonstrate how to organize data, calculate average revenue by genre, and create visualizations using the dplyr and ggplot2 packages. The goal is to show how these tools can help identify patterns in movie performance and make large datasets easier to understand.
In this tutorial, I will demonstrate how to import a movie dataset into R, explore the information it contains, and compare worldwide box office revenue across different movie genres. I will use the dplyr package to organize and summarize the data and ggplot2 to create a bar graph that makes the results easier to understand.
Being able to organize and visualize data is a useful skill in many different fields. When working with large datsets, it can be difficult to identify patterns just by looking at numbers. Using R makes it easier to summarize information and present findings in a way that others can understand. Although this tutorial focuses on movie genres and box office revenue, the same techniques can be applied to other types of data like customer behavior, sales, or program performance.
By the end of this code through, readers should be able to: 1. Import a real movie dataset into R and explore its contents. 2. Use the dplyr package to group movies by genre and calculate average worldwide box office revenue. 3. Create a bar graph using ggplot2 to compare revenue across different genres. 4. Interpret the results and explain how data visualization can help identify patterns.
Here, we’ll show how to use R to explore movie box office data and compare worldwide revenue across different genres. We’ll start by importing and examining the dataset, then use dplyr to organize the information and calculate average revenue. Finally, we’ll use ggplot2 to create a graph that helps visualize the difference between genres.
For this tutorial, I will be using a movie dataset from the TidyTuesday project, which provides real datasets for people to practice their data analysis skills. The dataset includes information such as movie titles, genres, production budgets, and worldwide box office revenue.
I will be using two R packages, dplyr and ggplot2. The dplyr package allows us to organize, group, and summarize data, while ggplot2 helps turn that information into graphs that are easier to interpret.
Since I work in movie theater operations, I thought it would be interesting to explore how different movie genres perform at the box office. This dataset gives us an opportunity to practice analyzing real-world information while learning R functions that can also be applied to other industries.
A basic example shows how to import and explore a movie dataset in R. Before analyzing the data, we need to load the packages that will help us organize and visualize the information. We will use dplyr to work with the dataset and ggplot2 to create graphs.
Next, we will use read.csv() to import the movie data from TidyTuesday. The head() function allows us to preview the first six rows, while names() shows the column names. These functions are useful because they help us understand how the dataset is organized before beginning our analysis.
# load packages
library(dplyr)
library(ggplot2)
# Import the movie dataset
movie_url <- "https://raw.githubusercontent.com/rfordatascience/tidytuesday/main/data/2018/2018-10-23/movie_profit.csv"
movies <- read.csv(movie_url)
# Preview the dataset
head(movies)## [1] "X" "release_date" "movie"
## [4] "production_budget" "domestic_gross" "worldwide_gross"
## [7] "distributor" "mpaa_rating" "genre"
More specifically, this can be used to compare worldwide box office revenue across different movie genres. First, we will remove movies with missing genre information or without a positive reported worldwide gross. Then, we will use dplyr to group movies by genre and calculate their average revenue.
The group_by() function organizes movies into categories based on their genre, while summarise() calculates the average revenue for each group. We will also count how many movies are included in each genre. Finally, arrange() will organize the results from highest to lowest average revenue.
# Calculate average worldwide revenue by genre
genre_summary <- movies %>%
filter(!is.na(genre), genre != "",
!is.na(worldwide_gross),
worldwide_gross > 0) %>%
group_by(genre) %>%
summarise(
movie_count = n(),
average_millions = mean(worldwide_gross) / 1000000,
.groups = "drop"
) %>%
arrange(desc(average_millions))
# View the results
print(genre_summary)## # A tibble: 5 × 3
## genre movie_count average_millions
## <chr> <int> <dbl>
## 1 Adventure 476 202.
## 2 Action 564 151.
## 3 Horror 295 67.9
## 4 Comedy 807 65.4
## 5 Drama 1223 53.8
Based on these results, Adventure movies have the highest
average worldwide box office revenue at approximately $202 million,
followed by Action movies at $151 million. Drama has the lowest average
revenue among the five genres at approximately $53.8 million, despite
having the largest number of movies in the dataset. This shows why
comparing averages can be useful instead of only looking at how many
movies belong to each genre.
What’s more, the information can also be used to create visualizations that make comparisons easier to understand. Although the previous table provides useful information, a graph allows us to see the differences between movie genres more clearly.
Using ggplot2, we will create a bar graph showing the average worldwide box office revenue for each genre. The geom_col() function creates the bars, while coord_flip() changes the direction of the graph to make the genre names easier to read. We will also add a title and axis labels to make the visualization easier to understand.
# Create a bar graph comparing movie genres
ggplot(genre_summary,
aes(x = reorder(genre, average_millions),
y = average_millions)) +
geom_col(fill = "steelblue") +
coord_flip() +
labs(
title = "Average Worldwide Box Office Revenue by Genre",
x = "Movie Genre",
y = "Average Worldwide Revenue (USD Millions)"
) +
theme_minimal()Most notably, these techniques are valuable for understanding how different statistics can affect the way we interpret data. Although the mean tells us the average worldwide revenue for each genre, it can be influenced by a few movies that earn significantly more than others.
To explore this further, we will compare the mean and median revenue for each genre. The median represents the middle value when the revenues are arranged from lowest to highest. Comparing these two measurements can help us determine whether a genre’s average is being influenced by a smaller number of highly successful movies.
This is important because relying only on averages may not always provide a complete picture of how movies perform.
Looking at the results, the mean revenue is higher than the median for every genre. For example, Adventure movies have an average revenue of approximately $202 million, while the median is only $123 million. This suggests that a smaller number of highly successful movies may be increasing the average. It also shows why looking at both the mean and median can provide a better understanding of movie performance.
# Compare mean and median revenue by genre
genre_comparison <- movies %>%
filter(!is.na(genre), genre != "",
!is.na(worldwide_gross),
worldwide_gross > 0) %>%
group_by(genre) %>%
summarise(
movies = n(),
mean_millions = round(
mean(worldwide_gross) / 1000000, 1
),
median_millions = round(
median(worldwide_gross) / 1000000, 1
),
.groups = "drop"
) %>%
arrange(desc(mean_millions))
# View the comparison
print(genre_comparison)## # A tibble: 5 × 4
## genre movies mean_millions median_millions
## <chr> <int> <dbl> <dbl>
## 1 Adventure 476 202 123.
## 2 Action 564 151. 88.5
## 3 Horror 295 67.9 36.7
## 4 Comedy 807 65.4 32.5
## 5 Drama 1223 53.8 21.4
To learn more about the dataset and R packages used in this tutorial, the following resources provide additional examples and explanations:
Resource I Introduction to dplyr
Resource II Introduction to ggplot2
Resource III TidyTuesday Movie Dataset
This code through references and cites the following sources:
R4DS Online Learning Community (2018). TidyTuesday Movie Profit Dataset. TidyTuesday GitHub Repository
Wickham et al. (2023). dplyr: A Grammar of Data Manipulation. dplyr Documentation
Wickham et al. (2025). ggplot2: Create Elegant Data Visualisations Using the Grammar of Graphics. ggplot2 Documentation