For this dataset, I will use state unemployment data from the U.S. Bureau of Labor Statistics. The original table reports unemployment rates for states across multiple time periods in separate columns, making it appropriate for a wide to long transformation. I will preserve the selected BLS table in its original wide structure as a CSV before performing any transformations. Using tidyr and dplyr, I will reshape the unemployment columns into a long format in which each row represents the unemployment rate for one state during one time period. The resulting dataset will contain variables such as state, period, and unemployment rate. I will standardize column names, convert values to appropriate data types, and investigate any missing or inconsistent observations. Using the tidy dataset, I will compare unemployment rates between the available periods and examine how unemployment changed across states. I plan to identify states with the largest increases and decreases, calculate summary statistics for unemployment rates, and compare individual states with the overall distribution. Visualizations will be used to display differences between states and changes over time.
Data Source
The data for this analysis comes from the U.S. Bureau of Labor Statistics (BLS) State Employment and Unemployment release. The dataset contains civilian labor force totals, numbers of unemployed individuals, and unemployment rates for U.S. states and selected areas across January 2025, November 2025, December 2025, and January 2026.
The original table was saved as unemployment_raw.csv and committed to GitHub before any transformations were performed.
The original dataset is stored in wide format. Each geographic area occupies one row, while measurements for different time periods are stored in separate columns. For example, unemployment rates for January 2025, November 2025, December 2025, and January 2026 are represented by four different columns.
This structure stores both the type of measurement and the time period within column names. To create a tidy dataset, these components will be separated into individual variables so that each row represents one state, measurement, and time period.
I use pivot_longer() to transform the repeated measurement columns into a long format. The original column names contain two pieces of information: the measurement being reported and the corresponding month and year. These components are separated into measure and period variables during the transformation.
The resulting dataset contains one observation for each geographic area, measurement type, and time period.
# A tibble: 12 × 4
state measure period value
<chr> <chr> <chr> <dbl>
1 Alabama Labor Force January 2025 2378428
2 Alabama Labor Force November 2025 2387873
3 Alabama Labor Force December 2025 2387810
4 Alabama Labor Force January 2026 2387649
5 Alabama Unemployed Number January 2025 72543
6 Alabama Unemployed Number November 2025 64802
7 Alabama Unemployed Number December 2025 64776
8 Alabama Unemployed Number January 2026 64061
9 Alabama Unemployment Rate January 2025 3.1
10 Alabama Unemployment Rate November 2025 2.7
11 Alabama Unemployment Rate December 2025 2.7
12 Alabama Unemployment Rate January 2026 2.7
The missing value check returned zero missing values across all four variables, so no imputation or row removal was necessary.
Analytical Methods
The analysis uses only the tidy version of the dataset. I focus on unemployment rates to compare geographic areas across the four reported periods and examine how unemployment changed between January 2025 and January 2026.
I calculate summary statistics for each period and identify the geographic areas with the largest increases and decreases in unemployment rates over the one-year period. Visualizations are used to show both overall unemployment patterns and changes across geographic areas.
The average unemployment rate across the 52 geographic areas increased from 3.89% in January 2025 to 4.13% in January 2026. The median increased from 3.85% to 4.30% over the same period. The maximum unemployment rate also increased from 5.7% to 6.7%, while the minimum increased from 2.0% to 2.2%.
The largest increase occurred in Delaware, where the unemployment rate rose from 4.1% to 5.4%, an increase of 1.3 percentage points. Minnesota, the District of Columbia, and Florida each increased by 1.0 percentage point, while Connecticut increased by 0.9 percentage points.
The largest decreases occurred in Indiana, Kentucky, and Ohio, where unemployment rates each fell by 0.5 percentage points. Alabama and Colorado each decreased by 0.4 percentage points. Overall, the results indicate that unemployment rates were somewhat higher across the included geographic areas in January 2026 than in January 2025, although the direction and magnitude of change varied considerably by location.
Conclusion
Transforming the BLS unemployment data from wide to long format separated the measurement type and reporting period that were originally embedded within the column names. The resulting tidy structure makes it possible to filter, group, summarize, and visualize labor-market measures consistently across geographic areas and time periods.
The analysis shows a modest overall increase in unemployment rates between January 2025 and January 2026. However, the changes were not uniform across geographic areas. Delaware experienced the largest increase, while Indiana, Kentucky, and Ohio experienced the largest decreases. This demonstrates how a tidy structure makes both overall trends and differences between individual geographic areas easier to identify.