Introduction

The objective of this project is to gain practical experience in transforming wide-format datasets into tidy datasets that are suitable for analysis and visualization. In many real-world datasets, related variables are stored across multiple columns, which can make the data more difficult to analyze. Following the tidy data principles described by Hadley Wickham, the data will be organized so that each variable forms a column, each observation forms a row, and each type of observational unit is represented appropriately. Using the tidyr and dplyr packages in R, this project will demonstrate how wide-format datasets can be transformed into tidy formats through reproducible data transformation workflows.


Each dataset will first be examined in its original format and then reshaped into a structure that is easier to analyze, summarize, and visualize. Three independent datasets from Discussion 5A will be used in this project:

  1. COVID-19 World Vaccination Progress
  2. Renewable Energy Capacity Time Series
  3. World GDP by Country (1960–2022)

Each dataset represents a different data transformation scenario. By restructuring these datasets into tidy formats, the project will demonstrate how appropriate data organization can support meaningful analysis and visualization.

Dataset 1: COVID-19 World Vaccination Progress

Data Source

Dataset: COVID World Vaccination Progress
Source: Kaggle
Link: https://www.kaggle.com/datasets/gpreda/covid-world-vaccination-progress

This dataset will be used to examine COVID-19 vaccination progress across countries. It contains several vaccination-related measures, including total vaccinations, the number of people vaccinated, the number of people fully vaccinated, and daily vaccination rates. The dataset will provide an opportunity to demonstrate how multiple vaccination measures can be reorganized into a tidy format. After the transformation, the data will be easier to compare across countries and analyze over time using R.

Structure Before Tidying


The dataset contains several vaccination-related measures stored in separate columns. Although these columns provide useful information, similar types of measurements are spread across multiple variables rather than being represented by a single variable that identifies the vaccination metric. This structure will be treated as a partially wide format because different vaccination measures are stored in separate columns. Planned Transformation.

The dataset will be reshaped by combining the vaccination metric columns into a single metric variable and placing their corresponding values into a value variable. This transformation will be performed using the pivot_longer() function from the tidyr package. The resulting tidy dataset will contain variables such as:

• country
• date
• metric
• value
This structure will make it easier to filter, group, compare, and visualize different vaccination metrics across countries and over time.

Planned Analysis

After the dataset is transformed into a tidy format, the analysis will focus on examining vaccination progress across countries and over time. Summary statistics and visualizations will be used to compare vaccination coverage and identify differences between countries. The analysis will include:

• Comparing vaccination coverage across countries
• Examining changes in vaccination progress over time
• Identifying countries with relatively higher or lower vaccination coverage
• Creating visualizations to highlight differences and trends in vaccination progress

Dataset 2: Renewable Energy Capacity Time Series

Data Source


Dataset: Renewable Power Plants / Renewable Capacity Time Series Source: Kaggle
Link: https://www.kaggle.com/datasets/eugeniyosetrov/renewable-power-plants

This dataset will be used to examine renewable energy generation across different countries and energy sources. It contains information related to several types of renewable energy, including solar, wind, hydro, and other renewable technologies. The dataset provides an opportunity to examine differences in renewable energy production across countries and over time.
## Structure Before Tidying
The dataset contains information organized across multiple columns, with different renewable energy sources and country-related information represented separately. Time-related information is also included to describe changes in energy generation over time. Because related information is distributed across multiple columns, the dataset will require restructuring before it can be easily analyzed. The wide format makes it more difficult to compare energy sources, countries, and time periods using standard R analysis and visualization functions.

Planned Transformation


The dataset will be transformed from a wide format into a longer, tidy format. The transformation will organize the data so that renewable energy type, country, time period, and generation value are represented as separate variables.


The transformation will use pivot_longer() from the tidyr package to combine related columns into appropriate variables. The resulting tidy dataset will make it easier to group and summarize the data by country, energy source, and time period.

The transformed dataset will include variables such as:
• country
• energy_type
• time
• generation

This structure will provide a consistent format for further analysis and visualization.

Planned Analysis


After the dataset is transformed into a tidy format, the analysis will focus on comparing renewable energy generation across countries, energy sources, and time periods.

The analysis will include:
• Comparing renewable energy generation across different energy sources
• Comparing renewable energy output between countries
• Examining changes in renewable energy generation over time
• Creating visualizations to identify differences and trends across countries and renewable energy technologies

These analyses will help demonstrate how tidy data can make it easier to compare renewable energy patterns and identify trends across different countries and energy sources.

Dataset 3: World GDP by Country (1960–2022)

Data Source


Dataset: World GDP by Country (1960–2022)
Source: Kaggle
Link: https://www.kaggle.com/datasets/annafabris/world-gdp-by-country-1960-2022

This dataset will be used to examine GDP trends across countries from 1960 through 2022. It contains GDP values for multiple countries, with annual values recorded for different years. The dataset will provide an opportunity to demonstrate how time-based data can be transformed from a wide format into a tidy structure.

Structure Before Tidying


The dataset stores each year as a separate column, creating a wide-format structure. Instead of having a single variable representing the year, the year values are used as column names. For example:

Country 1960 1961 1962 1963 … USA value value value value … France value value value value … Japan value value value value …

This structure makes it more difficult to analyze changes over time because the year information is stored in the column names rather than in a separate variable.

Planned Transformation

The dataset will be transformed from wide format into a tidy long format by combining the year columns into a single year variable and placing the corresponding GDP values into a gdp variable.
The transformation will be performed using the pivot_longer() function from the tidyr package. The resulting tidy dataset will contain variables such as:
• country
• year
• gdp

This structure will make it easier to filter, group, summarize, and visualize GDP data across countries and years.

Planned Analysis

After the dataset is transformed into a tidy format, the analysis will focus on examining GDP trends across countries and over time. The analysis will include:

• Comparing GDP trends across selected countries
• Examining changes in GDP over time
• Calculating and comparing GDP summaries across different periods
• Identifying countries with notable changes in GDP
• Creating time-series visualizations to illustrate economic trends

Reproducibility Plan

All data transformations and analyses will be implemented in a Quarto Markdown document using R and the tidyr, dplyr, and ggplot2 packages. The workflow for each dataset will include the following steps:
1. Import the original dataset from a CSV file.
2. Examine the structure and variables in the raw dataset.
3. Apply pivot_longer() and other dplyr functions to transform the data into a tidy format.
4. Analyze and summarize the transformed data.
5. Create tables and visualizations using ggplot2.

The complete workflow will be documented in the Quarto file so that the code and analysis can be reproduced from a clean R session.

Expected Outcome

The project will demonstrate how wide-format datasets from real-world sources can be transformed into tidy datasets that are easier to analyze and visualize. The transformed datasets will provide a structured format for comparing vaccination progress, examining renewable energy trends, and analyzing economic changes across countries. These transformations will also make it easier to summarize the data, create visualizations, and identify meaningful patterns and trends.

Conclusion

This project will demonstrate how wide-format datasets can be transformed into tidy datasets using R and the tidyr and dplyr packages. The three selected datasets COVID-19 vaccination progress, renewable energy data, and world GDP data have different structures, but each contains information that can be reorganized to make analysis easier.
Using transformations such as pivot_longer(), the datasets will be converted into formats where variables are represented as columns and observations as rows. This will make it easier to filter, summarize, compare, and visualize the data using dplyr and ggplot2.

Overall, this project will provide practical experience working with real-world datasets and applying tidy data principles. The final results will demonstrate how appropriate data transformations can make datasets easier to understand and prepare them for further analysis and visualization.