Introduction

Netflix has thousands of movies and TV shows from many countries. In this project I used R to explore a Netflix titles dataset. I wanted to find out what type of content Netflix has, which genres and countries are the biggest, how the content changed over the years, and how Indian content grew on Netflix.

Load data

url <- "https://raw.githubusercontent.com/rfordatascience/tidytuesday/master/data/2021/2021-04-20/netflix_titles.csv"
df <- read_csv(url)
glimpse(df)
## Rows: 7,787
## Columns: 12
## $ show_id      <chr> "s1", "s2", "s3", "s4", "s5", "s6", "s7", "s8", "s9", "s1…
## $ type         <chr> "TV Show", "Movie", "Movie", "Movie", "Movie", "TV Show",…
## $ title        <chr> "3%", "7:19", "23:59", "9", "21", "46", "122", "187", "70…
## $ director     <chr> NA, "Jorge Michel Grau", "Gilbert Chan", "Shane Acker", "…
## $ cast         <chr> "João Miguel, Bianca Comparato, Michel Gomes, Rodolfo Val…
## $ country      <chr> "Brazil", "Mexico", "Singapore", "United States", "United…
## $ date_added   <chr> "August 14, 2020", "December 23, 2016", "December 20, 201…
## $ release_year <dbl> 2020, 2016, 2011, 2009, 2008, 2016, 2019, 1997, 2019, 200…
## $ rating       <chr> "TV-MA", "TV-MA", "R", "PG-13", "PG-13", "TV-MA", "TV-MA"…
## $ duration     <chr> "4 Seasons", "93 min", "78 min", "80 min", "123 min", "1 …
## $ listed_in    <chr> "International TV Shows, TV Dramas, TV Sci-Fi & Fantasy",…
## $ description  <chr> "In a future where the elite inhabit an island paradise f…

Cleaning

colSums(is.na(df))
##      show_id         type        title     director         cast      country 
##            0            0            0         2389          718          507 
##   date_added release_year       rating     duration    listed_in  description 
##           10            0            7            0            0            0
df <- df %>%
  mutate(country = replace_na(country, "Unknown"),
         rating = replace_na(rating, "Not Rated"),
         date_added = mdy(date_added),
         year_added = year(date_added)) %>%
  distinct()

nrow(df)
## [1] 7787

I checked the missing values first. The country column had many blanks, so I filled them with “Unknown”. A few ratings were blank, so I filled them with “Not Rated”. The date_added column was text, so I changed it to date format and made a new year_added column. Then I removed duplicate rows.

Check my numbers (delete this section at the end)

df %>% count(type) %>% mutate(percent = round(n / sum(n) * 100, 1))
## # A tibble: 2 × 3
##   type        n percent
##   <chr>   <int>   <dbl>
## 1 Movie    5377    69.1
## 2 TV Show  2410    30.9
df %>% separate_rows(listed_in, sep = ", ") %>% count(listed_in, sort = TRUE) %>% head(5)
## # A tibble: 5 × 2
##   listed_in                  n
##   <chr>                  <int>
## 1 International Movies    2437
## 2 Dramas                  2106
## 3 Comedies                1471
## 4 International TV Shows  1199
## 5 Documentaries            786
df %>% count(country, sort = TRUE) %>% head(5)
## # A tibble: 5 × 2
##   country            n
##   <chr>          <int>
## 1 United States   2555
## 2 India            923
## 3 Unknown          507
## 4 United Kingdom   397
## 5 Japan            226
df %>% count(release_year, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
##   release_year     n
##          <dbl> <int>
## 1         2018  1121
## 2         2017  1012
## 3         2019   996
df %>% count(year_added, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
##   year_added     n
##        <dbl> <int>
## 1       2019  2153
## 2       2020  2009
## 3       2018  1685
df %>% count(rating, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
##   rating     n
##   <chr>  <int>
## 1 TV-MA   2863
## 2 TV-14   1931
## 3 TV-PG    806
df %>% filter(type == "Movie") %>%
  mutate(minutes = as.numeric(str_remove(duration, " min"))) %>%
  summarise(median = median(minutes, na.rm = TRUE),
            q25 = quantile(minutes, 0.25, na.rm = TRUE),
            q75 = quantile(minutes, 0.75, na.rm = TRUE))
## # A tibble: 1 × 3
##   median   q25   q75
##    <dbl> <dbl> <dbl>
## 1     98    86   114
df %>% filter(str_detect(country, "India")) %>% count(year_added, sort = TRUE) %>% head(3)
## # A tibble: 3 × 2
##   year_added     n
##        <dbl> <int>
## 1       2018   356
## 2       2019   241
## 3       2020   203

1. Movies vs TV Shows

ggplot(df, aes(x = type, fill = type)) +
  geom_bar(width = 0.6) +
  geom_text(stat = "count", aes(label = after_stat(count)), vjust = -0.5, size = 5) +
  scale_fill_manual(values = c("Movie" = "#E50914", "TV Show" = "#00B4D8")) +
  labs(title = "Movies vs TV Shows on Netflix", x = "", y = "Count") +
  my_theme + theme(legend.position = "none")

Netflix has more Movies than TV Shows. Movies are about 69% of all titles and TV Shows are about 31%.

2. Top 10 Genres

df %>%
  separate_rows(listed_in, sep = ", ") %>%
  count(listed_in, sort = TRUE) %>%
  slice_head(n = 10) %>%
  ggplot(aes(x = n, y = reorder(listed_in, n), fill = n)) +
  geom_col() +
  geom_text(aes(label = n), hjust = -0.2) +
  scale_fill_viridis_c(option = "plasma") +
  scale_x_continuous(expand = expansion(mult = c(0, 0.12))) +
  labs(title = "Top 10 Genres on Netflix", x = "Count", y = "") +
  my_theme + theme(legend.position = "none")

The most common genre is International Movies, followed by Dramas and Comedies. This shows Netflix focuses on international and drama content.

3. Top 10 Countries

df %>%
  count(country, sort = TRUE) %>%
  slice_head(n = 10) %>%
  ggplot(aes(x = n, y = reorder(country, n), fill = country)) +
  geom_col() +
  geom_text(aes(label = n), hjust = -0.2) +
  scale_fill_brewer(palette = "Spectral") +
  scale_x_continuous(expand = expansion(mult = c(0, 0.12))) +
  labs(title = "Top 10 Countries Producing Netflix Content", x = "Count", y = "") +
  my_theme + theme(legend.position = "none")

The United States produces the most titles, followed by India. “Unknown” appears because I filled the blank countries, and some rows list more than one country together.

4. Titles by Release Year

df %>%
  count(release_year) %>%
  ggplot(aes(x = release_year, y = n)) +
  geom_area(fill = "#E50914", alpha = 0.25) +
  geom_line(color = "#E50914", linewidth = 1.2) +
  labs(title = "Netflix Titles by Release Year", x = "Year", y = "Titles") +
  my_theme

Most titles were released after 2010, and the number rises sharply from about 2015. This shows Netflix added a lot of recent content.

5. Ratings

df %>%
  count(rating, sort = TRUE) %>%
  ggplot(aes(x = reorder(rating, -n), y = n, fill = rating)) +
  geom_col() +
  scale_fill_viridis_d(option = "plasma") +
  labs(title = "Content Ratings on Netflix", x = "Rating", y = "Count") +
  my_theme +
  theme(legend.position = "none", axis.text.x = element_text(angle = 45, hjust = 1))

The most common rating is TV-MA, which means most content is made for mature audiences. The second most common is TV-14.

6. Movie Duration

movies <- df %>%
  filter(type == "Movie") %>%
  mutate(minutes = as.numeric(str_remove(duration, " min")))

ggplot(movies, aes(x = minutes)) +
  geom_histogram(bins = 30, fill = "#00B4D8", color = "white") +
  geom_vline(aes(xintercept = median(minutes, na.rm = TRUE)),
             color = "#E50914", linewidth = 1.2, linetype = "dashed") +
  labs(title = "Movie Duration (minutes)",
       subtitle = "Red dashed line = median duration",
       x = "Minutes", y = "Count") +
  my_theme

Most movies are between about 87 and 114 minutes. The median is about 98 minutes.

7. My Own Question: Indian Titles Added per Year

df %>%
  filter(str_detect(country, "India")) %>%
  count(year_added) %>%
  drop_na() %>%
  ggplot(aes(x = year_added, y = n)) +
  geom_area(fill = "#FF9F1C", alpha = 0.3) +
  geom_line(color = "#FF9F1C", linewidth = 1.3) +
  geom_point(color = "#B20710", size = 3) +
  labs(title = "Indian Titles Added to Netflix per Year", x = "Year", y = "Titles") +
  my_theme

Indian titles added to Netflix rose over the years.

Key Findings

  1. About 69% of Netflix titles are Movies and 31% are TV Shows.
  2. The top genre is International Movies.
  3. The top country is the United States, followed by India.
  4. Most content was added to Netflix in the year.
  5. Indian content grew over the years, and the peak year was.

Conclusion

This analysis shows that Netflix has more Movies than TV Shows, with International Movies as the biggest genre and the United States as the top country. Netflix keeps adding more content over the years, and India is the second biggest country.


Project done by: Kayalvizhi Thiyagarajan

Course: BCA (Artificial Intelligence and Data Science)

Institute: Dr. M.G.R. Educational and Research Institute

Tools used: R, tidyverse, ggplot2, lubridate, R Markdown