Project 2 Data Tidying: Baby Names

Author

Zaina Hassan

Published

October 7, 2026

Approach

Dataset: SSA Top Baby Names

For this dataset, I will use the Social Security Administration’s historical table of the five most popular male and female baby names by year. The original data is presented in a wide format, with separate columns for each ranked female and male name. I will first recreate and save the original structure as a CSV file so that the raw data is preserved before any transformations are performed. Using tidyr and dplyr, I will transform the dataset from wide to long format so that each row represents a single name for a particular year, sex, and rank. Instead of maintaining separate columns for each ranking and sex, the tidy dataset will contain variables such as year, sex, rank, and name. I will also standardize the variable names and check for missing or inconsistent values. After tidying the data, I will analyze how the most popular baby names have changed over time. I plan to identify names that appeared in the top five most frequently, determine which names held the number-one position most often, and examine changes in naming patterns across years or decades. Visualizations will be used to show changes in name popularity and compare patterns between male and female names.

Data Source

The data for this analysis comes from the Social Security Administration (SSA) and contains the five most popular baby names for girls and boys for each year from 1926 through 2025. The original data was saved as baby_names_raw.csv and committed to GitHub before any transformations were performed.

The raw dataset is stored in a wide format, with separate columns for each combination of sex and rank.

Load Packages and Import Raw Data

Code
library(tidyverse)

baby_names_raw <- read.csv("baby_names_raw.csv")

head(baby_names_raw)
  Year Girls_Rank_1 Girls_Rank_2 Girls_Rank_3 Girls_Rank_4 Girls_Rank_5
1 2025       Olivia    Charlotte         Emma       Amelia       Sophia
2 2024       Olivia         Emma       Amelia    Charlotte          Mia
3 2023       Olivia         Emma    Charlotte       Amelia       Sophia
4 2022       Olivia         Emma    Charlotte       Amelia       Sophia
5 2021       Olivia         Emma    Charlotte       Amelia          Ava
6 2020       Olivia         Emma          Ava       Sophia    Charlotte
  Boys_Rank_1 Boys_Rank_2 Boys_Rank_3 Boys_Rank_4 Boys_Rank_5
1        Liam        Noah      Oliver    Theodore       Henry
2        Liam        Noah      Oliver    Theodore       James
3        Liam        Noah      Oliver       James      Elijah
4        Liam        Noah      Oliver       James      Elijah
5        Liam        Noah      Oliver      Elijah       James
6        Liam        Noah      Oliver      Elijah     William
Code
str(baby_names_raw)
'data.frame':   100 obs. of  11 variables:
 $ Year        : int  2025 2024 2023 2022 2021 2020 2019 2018 2017 2016 ...
 $ Girls_Rank_1: chr  "Olivia" "Olivia" "Olivia" "Olivia" ...
 $ Girls_Rank_2: chr  "Charlotte" "Emma" "Emma" "Emma" ...
 $ Girls_Rank_3: chr  "Emma" "Amelia" "Charlotte" "Charlotte" ...
 $ Girls_Rank_4: chr  "Amelia" "Charlotte" "Amelia" "Amelia" ...
 $ Girls_Rank_5: chr  "Sophia" "Mia" "Sophia" "Sophia" ...
 $ Boys_Rank_1 : chr  "Liam" "Liam" "Liam" "Liam" ...
 $ Boys_Rank_2 : chr  "Noah" "Noah" "Noah" "Noah" ...
 $ Boys_Rank_3 : chr  "Oliver" "Oliver" "Oliver" "Oliver" ...
 $ Boys_Rank_4 : chr  "Theodore" "Theodore" "James" "James" ...
 $ Boys_Rank_5 : chr  "Henry" "James" "Elijah" "Elijah" ...

Data Structure Before Tidying

In the original dataset, each row represents one year. However, the baby names are spread across ten separate columns: five ranked names for girls and five ranked names for boys. Therefore, information about sex and rank is stored within the column names rather than as separate variables.

This wide structure is useful for displaying rankings, but it is not ideal for analysis. A tidy structure should have one observation per row and separate variables for year, sex, rank, and name.

Code
dim(baby_names_raw)
[1] 100  11
Code
names(baby_names_raw)
 [1] "Year"         "Girls_Rank_1" "Girls_Rank_2" "Girls_Rank_3" "Girls_Rank_4"
 [6] "Girls_Rank_5" "Boys_Rank_1"  "Boys_Rank_2"  "Boys_Rank_3"  "Boys_Rank_4" 
[11] "Boys_Rank_5" 
Code
head(baby_names_raw)
  Year Girls_Rank_1 Girls_Rank_2 Girls_Rank_3 Girls_Rank_4 Girls_Rank_5
1 2025       Olivia    Charlotte         Emma       Amelia       Sophia
2 2024       Olivia         Emma       Amelia    Charlotte          Mia
3 2023       Olivia         Emma    Charlotte       Amelia       Sophia
4 2022       Olivia         Emma    Charlotte       Amelia       Sophia
5 2021       Olivia         Emma    Charlotte       Amelia          Ava
6 2020       Olivia         Emma          Ava       Sophia    Charlotte
  Boys_Rank_1 Boys_Rank_2 Boys_Rank_3 Boys_Rank_4 Boys_Rank_5
1        Liam        Noah      Oliver    Theodore       Henry
2        Liam        Noah      Oliver    Theodore       James
3        Liam        Noah      Oliver       James      Elijah
4        Liam        Noah      Oliver       James      Elijah
5        Liam        Noah      Oliver      Elijah       James
6        Liam        Noah      Oliver      Elijah     William

Transformation Steps

To convert the dataset from wide to tidy format, I use pivot_longer() to combine the ten ranking columns into two temporary variables: one containing the original column name and another containing the corresponding baby name.

I then use separate() to divide the original column names into sex and rank. Finally, I standardize the sex labels and convert rank to a numeric variable.

Code
baby_names_tidy <- baby_names_raw %>%
  pivot_longer(
    cols = -Year,
    names_to = "sex_rank",
    values_to = "name"
  ) %>%
  separate(
    sex_rank,
    into = c("sex", "rank"),
    sep = "_Rank_"
  ) %>%
  mutate(
    sex = recode(
      sex,
      "Girls" = "Female",
      "Boys" = "Male"
    ),
    rank = as.integer(rank)
  ) %>%
  rename(year = Year) %>%
  arrange(year, sex, rank)

head(baby_names_tidy, 10)
# A tibble: 10 × 4
    year sex     rank name    
   <int> <chr>  <int> <chr>   
 1  1926 Female     1 Mary    
 2  1926 Female     2 Dorothy 
 3  1926 Female     3 Betty   
 4  1926 Female     4 Helen   
 5  1926 Female     5 Margaret
 6  1926 Male       1 Robert  
 7  1926 Male       2 John    
 8  1926 Male       3 James   
 9  1926 Male       4 William 
10  1926 Male       5 Charles 

Check Results and Missing Values

Code
dim(baby_names_tidy)
[1] 1000    4
Code
str(baby_names_tidy)
tibble [1,000 × 4] (S3: tbl_df/tbl/data.frame)
 $ year: int [1:1000] 1926 1926 1926 1926 1926 1926 1926 1926 1926 1926 ...
 $ sex : chr [1:1000] "Female" "Female" "Female" "Female" ...
 $ rank: int [1:1000] 1 2 3 4 5 1 2 3 4 5 ...
 $ name: chr [1:1000] "Mary" "Dorothy" "Betty" "Helen" ...
Code
colSums(is.na(baby_names_tidy))
year  sex rank name 
   0    0    0    0 

After tidying, each row represents one ranked baby name for a specific year and sex. The final variables are year, sex, rank, and name. Missing values are checked after transformation so that they can be identified before beginning the analysis.

Analytical Methods

The analysis uses only the tidy version of the dataset. I examine which names appeared most frequently in the top five rankings, which names held the number-one position most often, and how the most popular names changed over time. I also compare patterns between female and male names. Counts and grouped summaries are calculated using dplyr, and visualizations are created using ggplot2.

Analysis 1: Most Frequent “Top 5” Names

Code
top_names <- baby_names_tidy %>%
  count(sex, name, sort = TRUE) %>%
  group_by(sex) %>%
  slice_max(order_by = n, n = 10, with_ties = FALSE) %>%
  ungroup()

top_names
# A tibble: 20 × 3
   sex    name            n
   <chr>  <chr>       <int>
 1 Female Mary           42
 2 Female Emma           24
 3 Female Linda          23
 4 Female Patricia       22
 5 Female Barbara        21
 6 Female Jennifer       21
 7 Female Jessica        21
 8 Female Olivia         21
 9 Female Sarah          19
10 Female Ashley         18
11 Male   James          62
12 Male   Michael        62
13 Male   John           47
14 Male   Robert         46
15 Male   David          39
16 Male   William        36
17 Male   Christopher    29
18 Male   Joshua         26
19 Male   Matthew        26
20 Male   Jacob          21

Analysis 2: Visualize the Most Frequent “Top 5” Names

Code
ggplot(top_names, aes(x = reorder(name, n), y = n, fill = sex)) +
  geom_col(show.legend = FALSE) +
  coord_flip() +
  facet_wrap(~sex, scales = "free_y") +
  scale_fill_manual(
    values = c("Female" = "lightpink", "Male" = "lightblue")
  ) +
  labs(
    title = "Most Frequent Top-Five Baby Names, 1926–2025",
    x = "Name",
    y = "Number of Years in the Top Five"
  ) +
  theme_minimal()

Analysis 3: Most Frequent #1 Names

Code
number_one_names <- baby_names_tidy %>%
  filter(rank == 1) %>%
  count(sex, name, sort = TRUE) %>%
  group_by(sex) %>%
  slice_max(order_by = n, n = 10, with_ties = FALSE) %>%
  ungroup()

number_one_names
# A tibble: 17 × 3
   sex    name         n
   <chr>  <chr>    <int>
 1 Female Mary        30
 2 Female Jennifer    15
 3 Female Emily       12
 4 Female Jessica      9
 5 Female Lisa         8
 6 Female Olivia       7
 7 Female Emma         6
 8 Female Linda        6
 9 Female Sophia       3
10 Female Ashley       2
11 Male   Michael     44
12 Male   Robert      15
13 Male   Jacob       14
14 Male   James       13
15 Male   Liam         9
16 Male   Noah         4
17 Male   David        1

Analysis 4: Time Based Visualization of Names

Code
number_one_by_year <- baby_names_tidy %>%
  filter(rank == 1)

ggplot(number_one_by_year, aes(x = year, y = name, fill = sex)) +
  geom_tile() +
  facet_wrap(~sex, scales = "free_y") +
  scale_fill_manual(
    values = c("Female" = "lightpink", "Male" = "lightblue")
  ) +
  labs(
    title = "Number-One Baby Names Over Time, 1926–2025",
    subtitle = "Each tile represents a year in which a name ranked #1",
    x = "Year",
    y = "Name"
  ) +
  guides(fill = "none") +
  theme_minimal()

Results

The analysis shows that several names remained popular for long periods of time. Among female names, Mary appeared in the top five for 42 of the 100 years, substantially more often than any other female name. Emma appeared for 24 years and Linda for 23 years. Among male names, James and Michael each appeared in the top five for 62 years, followed by John with 47 years and Robert with 46 years.

Looking specifically at the number-one ranked name reveals even stronger patterns. Mary was the most frequently ranked number-one female name, holding the top position for 30 years. Jennifer followed with 15 years, while Emily held the top position for 12 years. For male names, Michael was especially dominant, ranking number one for 44 years. Robert ranked first for 15 years, Jacob for 14 years, and James for 13 years.

The visualization of number-one names over time also shows that popularity often occurred in distinct periods. Female number-one names changed more frequently, while male rankings were dominated by a smaller number of names for longer periods. Michael’s extended period as the number-one male name is particularly visible in the data.

Code
number_one_summary <- number_one_names %>%
  rename(
    Sex = sex,
    Name = name,
    `Years Ranked #1` = n
  )

knitr::kable(
  number_one_summary,
  caption = "Most Frequent Number-One Baby Names, 1926–2025"
)
Most Frequent Number-One Baby Names, 1926–2025
Sex Name Years Ranked #1
Female Mary 30
Female Jennifer 15
Female Emily 12
Female Jessica 9
Female Lisa 8
Female Olivia 7
Female Emma 6
Female Linda 6
Female Sophia 3
Female Ashley 2
Male Michael 44
Male Robert 15
Male Jacob 14
Male James 13
Male Liam 9
Male Noah 4
Male David 1

Conclusion

Transforming the SSA baby-name data from wide to long format made it much easier to analyze changes in name popularity across years, sex, and rank. Instead of storing sex and rank information within separate column names, the tidy dataset represents these characteristics as individual variables that can be grouped, filtered, summarized, and visualized directly.

The analysis shows that some names maintained popularity for decades. Mary was the most persistent female name in both top-five appearances and number-one rankings, while James and Michael had the most top-five appearances among male names. Michael was particularly dominant at rank one, holding the position for 44 of the 100 years analyzed. Overall, the results also suggest greater turnover among number-one female names than among male names during this period.