str_glue("The longest title is {max_length} characters long.")
The longest title is 49 characters long.
str_length = determines how many characters are in the title
str_sub = it is picking off the characters 1-4 on the album year
str_to_lower = turn all the characters in the title lowercase
str_replace_all = replaces the dashes (-) with the slashes (/)
str_c = combines the different strings into one using the quotes
str_glue = same thing as str_c but uses the variable max_length without the quotes
Important functions for identifying strings which match
str_view() : most useful for testing str_subset() : useful for printing matches to the console str_detect() : useful when working within a tibble
Identify the input type and output type for each of these examples:
str_view(spot_smaller$subgenre, "pop")
[1] │ indie <pop>timism
[2] │ post-teen <pop>
[3] │ hip <pop>
[4] │ hip <pop>
[5] │ latin <pop>
typeof(str_view(spot_smaller$subgenre, "pop"))
[1] "character"
class(str_view(spot_smaller$subgenre, "pop"))
[1] "stringr_view"
str_view(spot_smaller$subgenre, "pop", match =NA)
[1] │ indie <pop>timism
[2] │ post-teen <pop>
[3] │ hip <pop>
[4] │ hip <pop>
[5] │ latin <pop>
[6] │ latin hip hop
[7] │ hip hop
[8] │ southern hip hop
[9] │ latin hip hop
[10] │ gangster rap
str_view(spot_smaller$subgenre, "pop", html =TRUE)
# A tibble: 5 × 4
title album_name artist subgenre
<chr> <chr> <chr> <chr>
1 Hear Me Now Hear Me Now Alok indie popti…
2 Run the World (Girls) 4 Beyoncé post-teen p…
3 Formation Lemonade Beyoncé hip pop
4 7/11 BEYONCÉ [Platinum Edition] Beyoncé hip pop
5 My Oh My (feat. DaBaby) Romance Camila Cabello latin pop
Find the mean song title length for songs with “pop” in the subgenre and songs without “pop” in the subgenre.
spot_smaller |>group_by(sub_pop =str_detect(subgenre, "pop")) |>mutate(sub_pop =ifelse(sub_pop ==FALSE, "Genre without pop", "Genre with pop")) |>summarize(mean_title_length =mean(str_length(title)))
# A tibble: 2 × 2
sub_pop mean_title_length
<chr> <dbl>
1 Genre with pop 13.6
2 Genre without pop 18.6
Producing a table like this would be great:
# A tibble: 2 × 2 sub_pop mean_title_length<lgl><dbl>1FALSE18.62TRUE13.6
Producing a table like this would be SUPER great (hint: ifelse()):
# A tibble: 2 × 2 sub_pop mean_title_length<chr><dbl>1 Genre with pop 13.62 Genre without pop 18.6
In the bigspotify dataset, find the proportion of songs which contain “love” in the title (track_name) by playlist_genre.
Here are the regular expression special characters that require an escape character (a preceding \): ^$.?*|+()[{
For any characters with special properties, use \ to “escape” its special meaning … but \ is itself a special character … so we need two \\! (e.g. \\$, \\., etc.)
str_view(spot_smaller$title, "$")
[1] │ Hear Me Now<>
[2] │ Run the World (Girls)<>
[3] │ Formation<>
[4] │ 7/11<>
[5] │ My Oh My (feat. DaBaby)<>
[6] │ It's Automatic<>
[7] │ Poetic Justice<>
[8] │ A.D.H.D<>
[9] │ Ya Estuvo<>
[10] │ Runnin (with A$AP Rocky, A$AP Ferg & Nicki Minaj)<>
In bigspotify, how many track_names include a $? Be sure you print the track_names you find and make sure the dollar sign is not just in a featured artist!
title <- bigspotify |>mutate(title_nofeat =str_remove(track_name, "\\(.*")) |>filter(str_detect(title_nofeat, "\\$")) |>distinct(track_name)title$track_name
[1] "Midnight Hour with Boys Noize & Ty Dolla $ign"
[2] "My First Kiss - feat. Ke$ha"
[3] "Wing$"
[4] "$Dreams"
[5] "$ave Dat Money (feat. Fetty Wap & Rich Homie Quan)"
[6] "NO TRU$T"
[7] "A$AP Forever"
[8] "M'$ (feat. Lil Wayne)"
[9] "Sie wollen meine Loui$ (Don Dollar)"
[10] "Foe Tha Love Of $"
[11] "A$AP"
[12] "$$$ - Remix"
[13] "Fre$h"
[14] "$ENHOR"
[15] "$20 Fine"
[16] "A$IAN BOY"
[17] "¿Cuánto E$?"
[18] "$. A. N. T. E. R. Í. A."
[19] "Bernice Burgo$"
[20] "$100 (feat. Polo Donatello)"
[21] "M'$"
[22] "Dat $tick"
[23] "Love$ick"
[24] "CA$H"
nrow(title)
[1] 24
In bigspotify, how many track_names include a dollar amount (a $ followed by a number).
Repetition
?
0 or 1 times
+
1 or more
*
0 or more
{n}
exactly n times
{n,}
n or more times
{,m}
at most m times
{n,m}
between n and m times
str_view(spot_smaller$album_name, "[A-Z]{2,}")
[4] │ <BEYONC>É [Platinum Edition]
[10] │ Creed <II>: The Album
Use at least 1 repetition symbol when solving 8-10 below
Modify the first regular expression above to also pick up “A.A” (in addition to “BEYONC” and “II”). That is, pick up strings where there might be a period between capital letters.
# A tibble: 10 × 2
title n_words
<chr> <int>
1 Hear Me Now 3
2 Run the World (Girls) 4
3 Formation 1
4 7/11 1
5 My Oh My (feat. DaBaby) 5
6 It's Automatic 2
7 Poetic Justice 2
8 A.D.H.D 1
9 Ya Estuvo 2
10 Runnin (with A$AP Rocky, A$AP Ferg & Nicki Minaj) 9
In the spot_smaller dataset, extract the first word from every title. Show how you would print out these words as a vector and how you would create a new column on the spot_smaller tibble. That is, produce this:
# A tibble: 10 × 2# title first_word# <chr> <chr> # 1 Hear Me Now Hear # 2 Run the World (Girls) Run # 3 Formation Formation # 4 7/11 7/11 # 5 My Oh My (feat. DaBaby) My # 6 It's Automatic It's # 7 Poetic Justice Poetic # 8 A.D.H.D A.D.H.D # 9 Ya Estuvo Ya #10 Runnin (with A$AP Rocky, A$AP Ferg & Nicki Minaj) Runnin spot_smaller |>select(title) |>mutate(first_word =str_extract(title, "^\\S+"))
# A tibble: 10 × 2
title first_word
<chr> <chr>
1 Hear Me Now Hear
2 Run the World (Girls) Run
3 Formation Formation
4 7/11 7/11
5 My Oh My (feat. DaBaby) My
6 It's Automatic It's
7 Poetic Justice Poetic
8 A.D.H.D A.D.H.D
9 Ya Estuvo Ya
10 Runnin (with A$AP Rocky, A$AP Ferg & Nicki Minaj) Runnin
Which decades are popular for playlist_names? Using the bigspotify dataset, try doing each of these steps one at a time!
filter the bigspotify dataset to only include playlists that include something like “80’s” or “00’s” in their title.
create a new column that extracts the decade
use count to find how many playlists include each decade
what if you include both “80’s” and “80s”?
how can you count “80’s” and “80s” together in your final tibble?
# A tibble: 12 × 1
playlist_name
<chr>
1 80's Songs | Top 💯 80s Music Hits
2 90's Southern Hip Hop
3 90's Hip Hop Ultimate Collection
4 Gangsta Rap/90's Hip-Hop
5 90's Gangster Rap
6 70's Classic Rock
7 Rock Ballads 80s 90s | Best Rock Love Songs 80's 90's Music Hits
8 2000's hard rock
9 80's Freestyle/Disco Dance Party (Set Crossfade to 4-Seconds)
10 90's NEW JACK SWING
11 New Jack Swing -late 80's & early 90's Hip Hop and R&B
12 R&B 80's/90's/00's
# A tibble: 12 × 2
playlist_name decade
<chr> <chr>
1 80's Songs | Top 💯 80s Music Hits 80's
2 90's Southern Hip Hop 90's
3 90's Hip Hop Ultimate Collection 90's
4 Gangsta Rap/90's Hip-Hop 90's
5 90's Gangster Rap 90's
6 70's Classic Rock 70's
7 Rock Ballads 80s 90s | Best Rock Love Songs 80's 90's Music Hits 80's
8 2000's hard rock 00's
9 80's Freestyle/Disco Dance Party (Set Crossfade to 4-Seconds) 80's
10 90's NEW JACK SWING 90's
11 New Jack Swing -late 80's & early 90's Hip Hop and R&B 80's
12 R&B 80's/90's/00's 80's
Describe to your groupmates what these expressions will match, and provide a word or expression as an example:
"(.)\\1\\1" This matches three of the same character in a row. The (.) finds one character, and \1\1 means that same character appears two more times.
"(.)(.)(.).*\\3\\2\\1" This matches a pattern where three characters appear at the beginning and then, after any number of characters, those same three characters appear in reverse order.
Construct a regular expression to match words in stringr::words that contain a repeated pair of letters (e.g. “church” contains “ch” repeated twice) but not match repeated pairs of numbers (e.g. 507-786-3861).
[1] "The canoe birch slid on the smooth planks."
[2] "Glue sheet the to the dark blue background."
[3] "It's to easy tell the depth of a well."
[4] "These a days chicken leg is a rare dish."
[5] "Rice often is served in round bowls."
The pattern takes the first three words of each sentence, and the replacement puts them back as word 1, word 3, word 2 — so it swaps the second and third words.