Netflix has thousands of movies and TV shows from many countries. In this project I used R to explore a Netflix titles dataset. I wanted to find out what type of content Netflix has, which genres and countries are the biggest, how the content changed over the years, and how Indian content grew on Netflix.
url <- "https://raw.githubusercontent.com/rfordatascience/tidytuesday/master/data/2021/2021-04-20/netflix_titles.csv"
df <- read_csv(url)
glimpse(df)
## Rows: 7,787
## Columns: 12
## $ show_id <chr> "s1", "s2", "s3", "s4", "s5", "s6", "s7", "s8", "s9", "s1…
## $ type <chr> "TV Show", "Movie", "Movie", "Movie", "Movie", "TV Show",…
## $ title <chr> "3%", "7:19", "23:59", "9", "21", "46", "122", "187", "70…
## $ director <chr> NA, "Jorge Michel Grau", "Gilbert Chan", "Shane Acker", "…
## $ cast <chr> "João Miguel, Bianca Comparato, Michel Gomes, Rodolfo Val…
## $ country <chr> "Brazil", "Mexico", "Singapore", "United States", "United…
## $ date_added <chr> "August 14, 2020", "December 23, 2016", "December 20, 201…
## $ release_year <dbl> 2020, 2016, 2011, 2009, 2008, 2016, 2019, 1997, 2019, 200…
## $ rating <chr> "TV-MA", "TV-MA", "R", "PG-13", "PG-13", "TV-MA", "TV-MA"…
## $ duration <chr> "4 Seasons", "93 min", "78 min", "80 min", "123 min", "1 …
## $ listed_in <chr> "International TV Shows, TV Dramas, TV Sci-Fi & Fantasy",…
## $ description <chr> "In a future where the elite inhabit an island paradise f…
colSums(is.na(df))
## show_id type title director cast country
## 0 0 0 2389 718 507
## date_added release_year rating duration listed_in description
## 10 0 7 0 0 0
df <- df %>%
mutate(country = replace_na(country, "Unknown"),
rating = replace_na(rating, "Not Rated"),
date_added = mdy(date_added),
year_added = year(date_added)) %>%
distinct()
nrow(df)
## [1] 7787
I checked the missing values first. The country column had many blanks, so I filled them with “Unknown”. A few ratings were blank, so I filled them with “Not Rated”. The date_added column was text, so I changed it to date format and made a new year_added column. Then I removed duplicate rows.
df %>% count(type) %>% mutate(percent = round(n / sum(n) * 100, 1))
## # A tibble: 2 × 3
## type n percent
## <chr> <int> <dbl>
## 1 Movie 5377 69.1
## 2 TV Show 2410 30.9
df %>% separate_rows(listed_in, sep = ", ") %>% count(listed_in, sort = TRUE) %>% head(5)
## # A tibble: 5 × 2
## listed_in n
## <chr> <int>
## 1 International Movies 2437
## 2 Dramas 2106
## 3 Comedies 1471
## 4 International TV Shows 1199
## 5 Documentaries 786
df %>% count(country, sort = TRUE) %>% head(5)
## # A tibble: 5 × 2
## country n
## <chr> <int>
## 1 United States 2555
## 2 India 923
## 3 Unknown 507
## 4 United Kingdom 397
## 5 Japan 226
df %>% count(release_year, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
## release_year n
## <dbl> <int>
## 1 2018 1121
## 2 2017 1012
## 3 2019 996
df %>% count(year_added, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
## year_added n
## <dbl> <int>
## 1 2019 2153
## 2 2020 2009
## 3 2018 1685
df %>% count(rating, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
## rating n
## <chr> <int>
## 1 TV-MA 2863
## 2 TV-14 1931
## 3 TV-PG 806
df %>% filter(type == "Movie") %>%
mutate(minutes = as.numeric(str_remove(duration, " min"))) %>%
summarise(median = median(minutes, na.rm = TRUE),
q25 = quantile(minutes, 0.25, na.rm = TRUE),
q75 = quantile(minutes, 0.75, na.rm = TRUE))
## # A tibble: 1 × 3
## median q25 q75
## <dbl> <dbl> <dbl>
## 1 98 86 114
df %>% filter(str_detect(country, "India")) %>% count(year_added, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
## year_added n
## <dbl> <int>
## 1 2018 356
## 2 2019 241
## 3 2020 203
ggplot(df, aes(x = type, fill = type)) +
geom_bar(width = 0.6) +
geom_text(stat = "count", aes(label = after_stat(count)), vjust = -0.5, size = 5) +
scale_fill_manual(values = c("Movie" = "#E50914", "TV Show" = "#00B4D8")) +
labs(title = "Movies vs TV Shows on Netflix", x = "", y = "Count") +
my_theme + theme(legend.position = "none")
Netflix has more Movies than TV Shows. Movies are about 69% of all titles and TV Shows are about 31%.
df %>%
separate_rows(listed_in, sep = ", ") %>%
count(listed_in, sort = TRUE) %>%
slice_head(n = 10) %>%
ggplot(aes(x = n, y = reorder(listed_in, n), fill = n)) +
geom_col() +
geom_text(aes(label = n), hjust = -0.2) +
scale_fill_viridis_c(option = "plasma") +
scale_x_continuous(expand = expansion(mult = c(0, 0.12))) +
labs(title = "Top 10 Genres on Netflix", x = "Count", y = "") +
my_theme + theme(legend.position = "none")
The most common genre is International Movies, followed by Dramas and Comedies. This shows Netflix focuses on international and drama content.
df %>%
count(country, sort = TRUE) %>%
slice_head(n = 10) %>%
ggplot(aes(x = n, y = reorder(country, n), fill = country)) +
geom_col() +
geom_text(aes(label = n), hjust = -0.2) +
scale_fill_brewer(palette = "Spectral") +
scale_x_continuous(expand = expansion(mult = c(0, 0.12))) +
labs(title = "Top 10 Countries Producing Netflix Content", x = "Count", y = "") +
my_theme + theme(legend.position = "none")
The United States produces the most titles, followed by India. “Unknown” appears because I filled the blank countries, and some rows list more than one country together.
df %>%
count(release_year) %>%
ggplot(aes(x = release_year, y = n)) +
geom_area(fill = "#E50914", alpha = 0.25) +
geom_line(color = "#E50914", linewidth = 1.2) +
labs(title = "Netflix Titles by Release Year", x = "Year", y = "Titles") +
my_theme
Most titles were released after 2010, and the number rises sharply from about 2015. This shows Netflix added a lot of recent content.
df %>%
count(rating, sort = TRUE) %>%
ggplot(aes(x = reorder(rating, -n), y = n, fill = rating)) +
geom_col() +
scale_fill_viridis_d(option = "plasma") +
labs(title = "Content Ratings on Netflix", x = "Rating", y = "Count") +
my_theme +
theme(legend.position = "none", axis.text.x = element_text(angle = 45, hjust = 1))
The most common rating is TV-MA, which means most content is made for mature audiences. The second most common is TV-14.
movies <- df %>%
filter(type == "Movie") %>%
mutate(minutes = as.numeric(str_remove(duration, " min")))
ggplot(movies, aes(x = minutes)) +
geom_histogram(bins = 30, fill = "#00B4D8", color = "white") +
geom_vline(aes(xintercept = median(minutes, na.rm = TRUE)),
color = "#E50914", linewidth = 1.2, linetype = "dashed") +
labs(title = "Movie Duration (minutes)",
subtitle = "Red dashed line = median duration",
x = "Minutes", y = "Count") +
my_theme
Most movies are between about 87 and 114 minutes. The median is about 98 minutes.
df %>%
filter(str_detect(country, "India")) %>%
count(year_added) %>%
drop_na() %>%
ggplot(aes(x = year_added, y = n)) +
geom_area(fill = "#FF9F1C", alpha = 0.3) +
geom_line(color = "#FF9F1C", linewidth = 1.3) +
geom_point(color = "#B20710", size = 3) +
labs(title = "Indian Titles Added to Netflix per Year", x = "Year", y = "Titles") +
my_theme
Indian titles added to Netflix rose over the years.
This analysis shows that Netflix has more Movies than TV Shows, with International Movies as the biggest genre and the United States as the top country. Netflix keeps adding more content over the years, and India is the second biggest country.
Project done by: Kayalvizhi Thiyagarajan
Course: BCA (Artificial Intelligence and Data Science)
Institute: Dr. M.G.R. Educational and Research Institute
Tools used: R, tidyverse, ggplot2, lubridate, R Markdown