project-2.1

Author

Tenzin Thakuri

1.Objective

This project develops practical competency in transforming wide-format datasets into tidy formats suitable for downstream analysis. All transformations are performed using the tidyr and dplyr packages in R.

2.Dataset Selection

Baby names dataset:https://www.ssa.gov/oact/babynames/top5names.html Obtained from 5A discussion post contributed by Aniss Sahraoui.

Data structure before tidying

This dataset is in wide structure format.Each row contains a year followed by ten names, with separate “Girls” and “Boys” sections that each repeat the headings “Rank 1” through “Rank 5.” Sex and rank are represented in the column headers, and multiple name-ranking observations appear in a single row.

3.Data Preparation

3.1 Raw Data Construction

library(readr)
library(dplyr)

Attaching package: 'dplyr'
The following objects are masked from 'package:stats':

    filter, lag
The following objects are masked from 'package:base':

    intersect, setdiff, setequal, union
library(tidyr)
library(naniar)
library(ggplot2)

baby_names_df<-read_csv("https://raw.githubusercontent.com/lhamo07/Data-607-Assignment/refs/heads/main/project-2.1/top5_baby_names.csv")
Rows: 100 Columns: 11
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (10): girl_rank1, girl_rank2, girl_rank3, girl_rank4, girl_rank5, boy_ra...
dbl  (1): year

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
baby_names_df
# A tibble: 100 × 11
    year girl_rank1 girl_rank2 girl_rank3 girl_rank4 girl_rank5 boy_rank1
   <dbl> <chr>      <chr>      <chr>      <chr>      <chr>      <chr>    
 1  2025 Olivia     Charlotte  Emma       Amelia     Sophia     Liam     
 2  2024 Olivia     Emma       Amelia     Charlotte  Mia        Liam     
 3  2023 Olivia     Emma       Charlotte  Amelia     Sophia     Liam     
 4  2022 Olivia     Emma       Charlotte  Amelia     Sophia     Liam     
 5  2021 Olivia     Emma       Charlotte  Amelia     Ava        Liam     
 6  2020 Olivia     Emma       Ava        Sophia     Charlotte  Liam     
 7  2019 Olivia     Emma       Ava        Sophia     Isabella   Liam     
 8  2018 Emma       Olivia     Ava        Isabella   Sophia     Liam     
 9  2017 Emma       Olivia     Ava        Isabella   Sophia     Liam     
10  2016 Emma       Olivia     Ava        Sophia     Isabella   Noah     
# ℹ 90 more rows
# ℹ 4 more variables: boy_rank2 <chr>, boy_rank3 <chr>, boy_rank4 <chr>,
#   boy_rank5 <chr>
miss_var_summary(baby_names_df)
# A tibble: 11 × 3
   variable   n_miss pct_miss
   <chr>       <int>    <num>
 1 year            0        0
 2 girl_rank1      0        0
 3 girl_rank2      0        0
 4 girl_rank3      0        0
 5 girl_rank4      0        0
 6 girl_rank5      0        0
 7 boy_rank1       0        0
 8 boy_rank2       0        0
 9 boy_rank3       0        0
10 boy_rank4       0        0
11 boy_rank5       0        0
glimpse(baby_names_df)
Rows: 100
Columns: 11
$ year       <dbl> 2025, 2024, 2023, 2022, 2021, 2020, 2019, 2018, 2017, 2016,…
$ girl_rank1 <chr> "Olivia", "Olivia", "Olivia", "Olivia", "Olivia", "Olivia",…
$ girl_rank2 <chr> "Charlotte", "Emma", "Emma", "Emma", "Emma", "Emma", "Emma"…
$ girl_rank3 <chr> "Emma", "Amelia", "Charlotte", "Charlotte", "Charlotte", "A…
$ girl_rank4 <chr> "Amelia", "Charlotte", "Amelia", "Amelia", "Amelia", "Sophi…
$ girl_rank5 <chr> "Sophia", "Mia", "Sophia", "Sophia", "Ava", "Charlotte", "I…
$ boy_rank1  <chr> "Liam", "Liam", "Liam", "Liam", "Liam", "Liam", "Liam", "Li…
$ boy_rank2  <chr> "Noah", "Noah", "Noah", "Noah", "Noah", "Noah", "Noah", "No…
$ boy_rank3  <chr> "Oliver", "Oliver", "Oliver", "Oliver", "Oliver", "Oliver",…
$ boy_rank4  <chr> "Theodore", "Theodore", "James", "James", "Elijah", "Elijah…
$ boy_rank5  <chr> "Henry", "James", "Elijah", "Elijah", "James", "William", "…

Rank are in char datatype, need to convert it to numeric datatype.

3.2 Data Import and Tidying

# Reshape to long format
baby_names_df_long <- baby_names_df %>%
# Rename variables to follow a consistent naming convention
  rename_with(tolower)%>%
  pivot_longer(
    cols = c(starts_with("girl_rank"), starts_with("boy_rank")),
    names_to = c("gender", "rank"),
    names_pattern = "(girl|boy)_rank(\\d)",
    values_to = "name"
  )%>%
# swap column rank and name
  relocate(name,.before=rank)%>%
  rename(birth_year=year)%>%
  # convert rank char to numeric
  mutate(rank=as.numeric(rank))
baby_names_df_long
# A tibble: 1,000 × 4
   birth_year gender name       rank
        <dbl> <chr>  <chr>     <dbl>
 1       2025 girl   Olivia        1
 2       2025 girl   Charlotte     2
 3       2025 girl   Emma          3
 4       2025 girl   Amelia        4
 5       2025 girl   Sophia        5
 6       2025 boy    Liam          1
 7       2025 boy    Noah          2
 8       2025 boy    Oliver        3
 9       2025 boy    Theodore      4
10       2025 boy    Henry         5
# ℹ 990 more rows
# Summarize missing data for all columns 
miss_var_summary(baby_names_df_long)
# A tibble: 4 × 3
  variable   n_miss pct_miss
  <chr>       <int>    <num>
1 birth_year      0        0
2 gender          0        0
3 name            0        0
4 rank            0        0

There is no any missing value in the dataset.

3.3 Analysis

# Average rank of each name
name_avg_rank <- baby_names_df_long %>%
  group_by(gender, name) %>%
  summarise(
    average_rank = mean(rank, na.rm = TRUE),
    years_in_top5 = n_distinct(birth_year),
  ) %>%
  arrange(average_rank)%>%
  ungroup()
`summarise()` has regrouped the output.
ℹ Summaries were computed grouped by gender and name.
ℹ Output is grouped by gender.
ℹ Use `summarise(.groups = "drop_last")` to silence this message.
ℹ Use `summarise(.by = c(gender, name))` for per-operation grouping
  (`?dplyr::dplyr_by`) instead.
name_avg_rank
# A tibble: 73 × 4
   gender name     average_rank years_in_top5
   <chr>  <chr>           <dbl>         <int>
 1 girl   Mary             1.36            42
 2 boy    Liam             1.38            13
 3 boy    Michael          1.48            62
 4 girl   Emily            1.62            16
 5 boy    Jacob            1.67            21
 6 girl   Lisa             1.77            13
 7 girl   Jennifer         1.81            21
 8 girl   Jessica          1.86            21
 9 girl   Emma             2               24
10 boy    Noah             2.07            15
# ℹ 63 more rows

Mary was one of the most popular girls’ names, with an average rank of 1.36 over 42 years Liam had a slightly better average rank than Michael, but Michael stayed in the top five for much longer—62 years compared with Liam’s 13 years.

name_avg_rank %>%
  filter(gender == "girl", years_in_top5 >= 10) %>%
  slice_min(average_rank, n = 10, with_ties = FALSE) %>%
  ggplot(aes(
    x = reorder(name, average_rank),
    y = average_rank,
    fill=average_rank
  )) +
  geom_col() +
  labs(
    title = "Girls' Names with the Best Average Rankings",
    x = "Baby Name",
    y = "Average Rank",
  fill = "Average Rank"

  ) +
  theme_minimal()

Marry is the Top-ranked name and Barbara is the lowest-ranked name.

name_avg_rank %>%
  filter(gender == "boy", years_in_top5 >= 10) %>%
  slice_min(average_rank, n = 10, with_ties = FALSE) %>%
  ggplot(aes(
    x = reorder(name, average_rank),
    y = average_rank,
    fill=average_rank
    
  )) +
  geom_col() +
  labs(
    title = "Boys' Names with the Best Average Rankings",
    x = "Baby Name",
    y = "Average Rank",
    fill = "Average Rank"

    
    
  ) +
  theme_minimal()

Liam is the Top-ranked name and John is the lowest-ranked name.

4.Conclusion

This dataset helped me transform the baby names dataset from wide format to tidy long format using dplyr. By analyzing average rankings and the number of years names appeared in the top five, Iwas able to identified popular boy and girl names in the year 1926-2025.

5.Citation

Anthropic. (2026). Claude Sonnet 5.5 [Large language model]. https://claude.ai. Accessed October 11, https://claude.ai/share/0b74850c-e792-459a-899a-89eeb306cae7