Lets load R packages to start our activity

library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.1.4     ✔ readr     2.1.5
## ✔ forcats   1.0.0     ✔ stringr   1.5.1
## ✔ ggplot2   3.5.1     ✔ tibble    3.2.1
## ✔ lubridate 1.9.3     ✔ tidyr     1.3.1
## ✔ purrr     1.0.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(skimr)
library(tidytext)
library(showtext)
## Loading required package: sysfonts
## Loading required package: showtextdb
library(flextable)
## 
## Attaching package: 'flextable'
## 
## The following object is masked from 'package:purrr':
## 
##     compose

To begin, import the swiftSongs.csv file located here using the code below:

# Variables to keep
keeps <- c("track_name", "album_name", "youtube_title", "youtube_duration", "full_lyrics")

# Importing CSV file
swift_songs <- read_csv("https://raw.githubusercontent.com/dilernia/STA418-518/main/Data/swiftSongs.csv") |> dplyr::select(all_of(keeps))
## Rows: 151 Columns: 34
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr  (10): youtube_title, youtube_description, youtube_duration, youtube_url...
## dbl  (22): youtube_view_count, youtube_like_count, youtube_favorite_count, y...
## lgl   (1): explicit
## dttm  (1): youtube_publish_date
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Lets conduct some explotary data analysis

glimpse(swift_songs)
## Rows: 151
## Columns: 5
## $ track_name       <chr> "...Ready For It?", "‘tis the damn season", "august",…
## $ album_name       <chr> "reputation", "evermore", "folklore", "folklore", "fo…
## $ youtube_title    <chr> "Taylor Swift - …Ready For It?", "Taylor Swift - ‘tis…
## $ youtube_duration <chr> "PT3M31S", "PT3M56S", "PT4M24S", "PT4M56S", "PT4M35S"…
## $ full_lyrics      <chr> "Knew he was a killer first time that I saw him Wonde…
skim(swift_songs)
Data summary
Name swift_songs
Number of rows 151
Number of columns 5
_______________________
Column type frequency:
character 5
________________________
Group variables None

Variable type: character

skim_variable n_missing complete_rate min max empty n_unique whitespace
track_name 0 1 2 70 0 151 0
album_name 0 1 3 12 0 10 0
youtube_title 0 1 5 79 0 151 0
youtube_duration 0 1 4 7 0 92 0
full_lyrics 0 1 786 3505 0 151 0

Matching Strings

# Detecting if a string contains the substring 'Taylor'
str_detect(string = c("Taylor Swift", "Taylor Lautner", "Harry Styles"),
           pattern = "Taylor")
## [1]  TRUE  TRUE FALSE

Using the str_detect() and mutate() functions, add a new boolean variable called contains_midnight to swift_songs that indicates whether or not a song’s lyrics contain the word “midnight”.

swift_songs<-swift_songs |>
  mutate(contains_midnight = str_detect(full_lyrics,pattern = "midnight"))

How many of Taylor’s songs mention the word “midnight”?

swift_songs |>
  count(contains_midnight) |>
  flextable()

contains_midnight

n

FALSE

145

TRUE

6

How many of Taylor’s songs mention the word “midnight” or “Midnight”?

swift_songs <- swift_songs |>
  mutate(contains_midnight = str_detect(string = str_to_lower(full_lyrics), 
                                        pattern = "midnight|Midnight"))


swift_songs |>
  count(contains_midnight) |>
  flextable()

contains_midnight

n

FALSE

143

TRUE

8

# Counting instances of "man"
str_count("I’m so sick of running as fast as I can Wondering if I'd get there quicker if I was a man And I'm so sick of them coming at me again 'Cause if I was a man, then I'd be the man I'd be the man I'd be the man", 
          pattern = "man")
## [1] 5

You try

Using the str_count() and mutate() functions, add a new variable called love_count to swift_songs that indicates how many times each song mentions the word “love” agnostic of case.

swift_songs <- swift_songs |> mutate(love_count = str_count(string = str_to_lower(full_lyrics), pattern="love"))
swift_songs
## # A tibble: 151 × 7
##    track_name              album_name youtube_title youtube_duration full_lyrics
##    <chr>                   <chr>      <chr>         <chr>            <chr>      
##  1 ...Ready For It?        reputation Taylor Swift… PT3M31S          "Knew he w…
##  2 ‘tis the damn season    evermore   Taylor Swift… PT3M56S          "If I want…
##  3 august                  folklore   Taylor Swift… PT4M24S          "Salt air,…
##  4 betty                   folklore   Taylor Swift… PT4M56S          "Betty, I …
##  5 cardigan                folklore   Taylor Swift… PT4M35S          "Vintage t…
##  6 champagne problems      evermore   Taylor Swift… PT4M8S           "You booke…
##  7 closure                 evermore   Taylor Swift… PT3M3S           "It's been…
##  8 coney island (feat. Th… evermore   Taylor Swift… PT4M38S          "Break my …
##  9 cowboy like me          evermore   Taylor Swift… PT4M40S          "And the t…
## 10 dorothea                evermore   Taylor Swift… PT3M45S          "Hey, Doro…
## # ℹ 141 more rows
## # ℹ 2 more variables: contains_midnight <lgl>, love_count <int>

Which song mentions love the most times, and how many times is it mentioned?

swift_songs |>
  arrange(desc(love_count)) |>
  dplyr::select(-full_lyrics) |>
  slice(1:1) |>
  flextable()

track_name

album_name

youtube_title

youtube_duration

contains_midnight

love_count

This Love

1989

This Love

PT4M11S

FALSE

52

sum(swift_songs$love_count)
## [1] 331

Modifying strings

There are also several functions for mutating or modifying character strings. For example, the str_c() function concatenates or combines multiple strings together.

str_c("...Are you ", "ready for it?")
## [1] "...Are you ready for it?"
str_c(letters, LETTERS)
##  [1] "aA" "bB" "cC" "dD" "eE" "fF" "gG" "hH" "iI" "jJ" "kK" "lL" "mM" "nN" "oO"
## [16] "pP" "qQ" "rR" "sS" "tT" "uU" "vV" "wW" "xX" "yY" "zZ"
# Positive indices start from beginning of the string
str_sub(string = "Eras Tour", start = 1, end = 3)
## [1] "Era"
# Negative indices start from end of the string
str_sub(string = "Eras Tour", start = -4, end = -1)
## [1] "Tour"
str_replace_all("I’m so sick of running as fast as I can Wondering if I'd get there quicker if I was a man And I'm so sick of them coming at me again 'Cause if I was a man, then I'd be the man I'd be the man I'd be the man", 
                pattern = "man", replacement = "!!!")
## [1] "I’m so sick of running as fast as I can Wondering if I'd get there quicker if I was a !!! And I'm so sick of them coming at me again 'Cause if I was a !!!, then I'd be the !!! I'd be the !!! I'd be the !!!"

You try

Create a new variable called youtube_time that is the same as youtube_duration, but with a : symbol replacing the M.

swift_songs <- swift_songs |> mutate(youtube_time = str_replace_all(youtube_duration, pattern="m|M", replacement = ":"),
                                     youtube_time = str_remove_all(youtube_time, pattern="PT|S"),
                                     youtube_time = case_when(str_length(youtube_time) == 2 ~ str_replace_all(youtube_time, pattern=":", replacement=":00"),
          str_length(youtube_time) == 3 ~ str_replace_all(youtube_time, pattern=":", replacement=":0"),
                                                     TRUE ~ youtube_time))

swift_songs$youtube_time
##   [1] "3:31" "3:56" "4:24" "4:56" "4:35" "4:08" "3:03" "4:38" "4:40" "3:45"
##  [11] "4:51" "5:05" "4:47" "3:09" "5:18" "3:42" "3:12" "4:14" "4:19" "3:41"
##  [21] "3:59" "4:18" "3:30" "4:17" "3:38" "3:55" "3:30" "3:32" "3:52" "3:16"
##  [31] "4:08" "4:13" "4:03" "3:41" "3:20" "3:45" "5:30" "3:14" "5:10" "5:04"
##  [41] "3:32" "4:11" "5:54" "3:38" "4:33" "4:26" "3:27" "4:33" "4:31" "4:00"
##  [51] "3:43" "4:49" "3:00" "3:32" "4:55" "6:44" "3:20" "3:55" "3:57" "3:51"
##  [61] "5:53" "4:12" "4:12" "3:22" "4:11" "5:04" "3:46" "3:54" "3:41" "3:32"
##  [71] "4:03" "4:15" "3:22" "4:08" "4:03" "3:59" "2:52" "5:58" "3:16" "2:55"
##  [81] "3:28" "3:39" "5:03" "3:24" "2:32" "3:30" "3:35" "4:12" "6:08" "3:27"
##  [91] "3:12" "5:18" "4:16" "3:57" "3:59" "3:43" "3:34" "3:15" "4:09" "4:04"
## [101] "3:00" "3:56" "3:56" "4:51" "3:56" "3:23" "4:17" "3:44" "3:32" "3:35"
## [111] "4:02" "4:44" "4:02" "4:03" "4:20" "3:48" "3:23" "4:31" "4:01" "3:38"
## [121] "4:56" "3:57" "3:25" "4:03" "3:13" "3:38" "3:00" "3:21" "3:40" "4:36"
## [131] "3:48" "4:01" "4:15" "4:46" "3:28" "4:27" "4:05" "3:28" "4:11" "4:09"
## [141] "3:52" "4:01" "2:50" "3:36" "3:33" "4:05" "3:55" "3:49" "3:31" "4:23"
## [151] "3:19"

Modify youtube_time by removing the P, T, and S letters. Notice that the seconds component is not always correct in our newly created youtube_time variable. We can fix this using the case_when() function together with str_replace_all() and str_length() which returns the number of characters in a string

# Modify youtube_time by removing P, T, and S letters.
swift_songs <- swift_songs %>%
  mutate(youtube_time  = str_replace_all(youtube_duration,
                                         pattern = "M",
                                         replacement = ":"),
         youtube_time = str_remove_all(youtube_time, pattern = "PT|S"))

Modify youtube_time to add trailing or extra 0’s when needed using the case_when() function together with str_replace_all() and str_length().

swift_songs <- swift_songs %>%
  mutate(youtube_time = case_when(str_length(youtube_time) == 2 ~ str_c(youtube_time, "00"),
                                  str_length(youtube_time) == 3 ~ str_replace_all(
                                                                          youtube_time,
                                                                          pattern = ":",
                                                                          replacement = ":0"),
                                             TRUE ~ youtube_time))

Coerce youtube_time to be a special date / time variable using the parse_date_time() function from the lubridate package using the code below.

# Coercing youtube_time to a date / time variable
swift_songs <- swift_songs |> 
  dplyr::mutate(youtube_time = lubridate::parse_date_time(youtube_time, orders = "%M:%S"))

Use the minute() and second() functions from the lubridate package, create a new variable, song_duration_s that gives the song duration in seconds using the code below.

# Creating song_duration_s variable
swift_songs <- swift_songs |> 
  dplyr::mutate(song_duration_s = lubridate::second(youtube_time) + 
                               60*lubridate::minute(youtube_time))

The escape sequence \w+ can be used to match any ‘word’ character (although it very slightly over counts). Create a new variable song_words equal to the number of words in the song using the str_count() function and the full_lyrics variable.

# Creating song_words variable
swift_songs <- swift_songs |> 
  dplyr::mutate(song_words = str_count(full_lyrics, pattern = "\\w+"))

Reproduce the plot below showing the relationship between the duration of each song in seconds and its number of words. Hint: to match the style of the points, use fill = ‘#01a7d9’, pch = 23, color = ‘#7d488e’ inside of the geom_point() layer.

#creating a scatter plot 
swift_songs |>
  ggplot(aes(x = song_words, 
             y  = song_duration_s)) +
  
  geom_point(fill = '#01a7d9', pch = 23, color = '#7d488e') +
  
  labs (title = "Number of words by Taylor Swift song duration", 
        x = "Number of words in lyrics", 
        t = "Song Duration(seconds)", 
        caption = "Data source: geniusr R package") +
  theme_bw()

Capitalization and spacing

# setting all characters to lowercase
str_to_lower("It's nice to have a friend")
## [1] "it's nice to have a friend"
# setting all characters to uppercase
str_to_lower("It's nice to have a friend")
## [1] "it's nice to have a friend"
# setting all characters to title case
str_to_lower("It's nice to have a friend")
## [1] "it's nice to have a friend"
# Removing spaces at start and end of string
str_trim(" Best believe I'm still bejeweled     When I walk in the room     I can still make the whole place shimmer ")
## [1] "Best believe I'm still bejeweled     When I walk in the room     I can still make the whole place shimmer"
# Removing spaces at start and end of string and repetitive spaces
str_squish(" Best believe I'm still bejeweled     When I walk in the room     I can still make the whole place shimmer ")
## [1] "Best believe I'm still bejeweled When I walk in the room I can still make the whole place shimmer"

Making a ☁️Word Cloud☁️

library(wordcloud2) 

# Tallying up frequency of words in all songs
wordFreqs <- swift_songs |> 
  unnest_tokens(word, full_lyrics) |> 
  count(word) 

# Removing 'stop words' (common but not very meaningful words)
wordFreqs <- wordFreqs |> 
  anti_join(stop_words) |>
  filter(!(word %in% c("ooh", "yeah", "a.m")))
## Joining with `by = join_by(word)`

Finally, we can create the word cloud using the wordcloud2() function from the wordcloud2 package.

wordcloud2(wordFreqs, size=1.6, color='random-dark')

We can also customize the font using Google fonts and customize the colors used in the plot. For more customization options using wordcloud2, see the Wordcloud2 library. Note that custom fonts can be finicky in R, so it is okay if you cannot get this part to work for the activity.

library(showtext) # for custom fonts

font_family <- "satisfy"

word_colors <- c('#7f6070', '#964c32', '#bb9559',
                  '#8c8c8c', '#eeadcf', '#7193ac',
                  '#a81e47', '#0c0c0c', '#7d488e', '#01a7d9')

# Downloading Google fonts for plots
font_add_google(name = str_to_title(font_family), 
                family = font_family)

# Creating word cloud
set.seed(1989)

my_word_cloud <- wordcloud2(wordFreqs, size = 1.6, 
                          color = sample(word_colors, 
                                         replace = TRUE, size = nrow(wordFreqs)), 
                          fontFamily = font_family)

my_word_cloud

Sentiment analysis

 # Tokenizing song lyrics into words for each song
swift_tidy <- swift_songs |> 
  unnest_tokens(word, full_lyrics)
# Getting sentiments
bing_sentiments <- get_sentiments("bing")
bing_sentiments |>
  slice_head(n = 4)
## # A tibble: 4 × 2
##   word       sentiment
##   <chr>      <chr>    
## 1 2-faces    negative 
## 2 abnormal   negative 
## 3 abolish    negative 
## 4 abominable negative

Using the bing sentiments lexicon, let’s explore the sentiment of songs from Taylor’s critically acclaimed 1989 album.

library(dplyr)
# Summarizing sentiment of each song from 1989
swift_1989_sentiment <- swift_tidy |>
  dplyr::filter(album_name == "1989") |> 
  dplyr::select(track_name, album_name, song_words, word) |> 
  inner_join(bing_sentiments) |>
  group_by(track_name) |> 
  dplyr::mutate(word_num = 1:n()) |> 
  ungroup() 
## Joining with `by = join_by(word)`

Use the vector below to reproduce the plot visualizing the sentiment for songs from 1989, and use a scale_fill_manual() layer with the colors c(‘#7193ac’, ‘#01a7d9’) to match the coloring as well.

# Define a vector with the 1989 songs in order

songs_1989 <- c("Welcome To New York", "Blank Space", "Style", "Out Of The Woods",
                     "All You Had To Do Was Stay", "Shake It Off", "I Wish You Would",
                     "Bad Blood", "Wildest Dreams", "How You Get The Girl",
                     "This Love", "I Know Places", "Clean")

# Creating wide version of data to plot net sentiment
swift_1989_sentiment_wide <- swift_1989_sentiment |> 
  count(track_name, index = word_num, sentiment) |>
  pivot_wider(names_from = sentiment, values_from = n, values_fill = 0) |>
  mutate(sentiment = positive - negative) |>
  mutate(sentiment = positive - negative, 
         track_name = fct_relevel(track_name, songs_1989))
#plotting the sentiment for each song 
swift_1989_sentiment_wide |>
 
  ggplot(aes(x = index, 
             y = sentiment, 
             fill = sentiment > 0)) +
  geom_col() +
  scale_fill_manual(values = c("#7193ac", "#01a7d9"))+
  scale_y_continuous(breaks = c(-1, 0, 1)) +
 facet_wrap(~ track_name, ncol = 4, scales = "free_x") +
  labs(title = " Sentimment of lyrics from Taylor Swift's 1989 Album ", 
        x = "", 
        t = "Sentiment of each word", 
        caption = "Data source: geniusr R package \n Sentiment calculated using bing sentiment lexicon")+
  ggthemes::theme_few() +
  theme(legend.position = "none", 
        axis.ticks.x = element_blank(), 
        axis.text.x = element_blank())

What are some limitations with this analysis?

Sentiment analysis is limited by subjectivity, context sensitivity, and difficulty in interpreting linguistic nuances, such as negations and sarcasm. Consider the phrase “I love this movie, not!” In this example, the sentiment analysis algorithm might initially interpret “love” as positive, but the presence of “not” reverses the sentiment, indicating sarcasm or negative sentiment, illustrating the challenge in accurately capturing sentiment nuances.

Bonus(Optional)

Using the title of the YouTube video for each song, create a variable indicating whether or not the video is an official music video, official lyric video, or other type of video.

# Create a new variable to store the type of video
swift_songs <- swift_songs |>
  mutate(video_type = case_when(
    str_detect(youtube_title, regex("official music video", ignore_case = TRUE)) ~ "Official Music Video",
    str_detect(youtube_title, regex("official lyric video", ignore_case = TRUE)) ~ "Official Lyric Video",
    TRUE ~ "Other"
  ))

# View the dataset with the new variable
swift_songs
## # A tibble: 151 × 11
##    track_name              album_name youtube_title youtube_duration full_lyrics
##    <chr>                   <chr>      <chr>         <chr>            <chr>      
##  1 ...Ready For It?        reputation Taylor Swift… PT3M31S          "Knew he w…
##  2 ‘tis the damn season    evermore   Taylor Swift… PT3M56S          "If I want…
##  3 august                  folklore   Taylor Swift… PT4M24S          "Salt air,…
##  4 betty                   folklore   Taylor Swift… PT4M56S          "Betty, I …
##  5 cardigan                folklore   Taylor Swift… PT4M35S          "Vintage t…
##  6 champagne problems      evermore   Taylor Swift… PT4M8S           "You booke…
##  7 closure                 evermore   Taylor Swift… PT3M3S           "It's been…
##  8 coney island (feat. Th… evermore   Taylor Swift… PT4M38S          "Break my …
##  9 cowboy like me          evermore   Taylor Swift… PT4M40S          "And the t…
## 10 dorothea                evermore   Taylor Swift… PT3M45S          "Hey, Doro…
## # ℹ 141 more rows
## # ℹ 6 more variables: contains_midnight <lgl>, love_count <int>,
## #   youtube_time <dttm>, song_duration_s <dbl>, song_words <int>,
## #   video_type <chr>

Use the str_glue() function in tandem with ggplot to create a scatter plot showing the relationship between the total number of characters in each song’s lyrics (full_lyrics) and the total number of characters in each song’s title (track_name), including the correlation between the two variables dynamically in the subtitle.

# creating 'video_type' variable
swift_songs <- swift_songs %>%
  mutate(video_type = case_when(
    str_detect(youtube_title, regex("official music video", ignore_case = TRUE)) ~ "Official Music Video",
    str_detect(youtube_title, regex("official lyric video", ignore_case = TRUE)) ~ "Official Lyric Video",
    TRUE ~ "Other"
  ))

# Calculate total number of characters in song title and lyrics
swift_songs$total_title_chars <- str_length(swift_songs$track_name)
swift_songs$total_lyrics_chars <- str_length(swift_songs$full_lyrics)

# Create scatter plot with colors based on video_type
ggplot(swift_songs, aes(x = total_lyrics_chars, y = total_title_chars, color = video_type)) +
  geom_point() +
  labs(
    x = "Total number of characters in lyrics",
    y = "Total number of characters in title",
    title = "Relationship between Title and Lyrics Length",
    subtitle = str_glue("Correlation: {cor(swift_songs$total_lyrics_chars, swift_songs$total_title_chars)}")
  )