The objective of this project is to gain practical experience in transforming wide-format datasets into tidy datasets that are suitable for analysis and visualization. In many real-world datasets, related variables are stored across multiple columns, which can make the data more difficult to analyze. Following the tidy data principles described by Hadley Wickham, the data will be organized so that each variable forms a column, each observation forms a row, and each type of observational unit is represented appropriately. Using the tidyr and dplyr packages in R, this project will demonstrate how wide-format datasets can be transformed into tidy formats through reproducible data transformation workflows.
Each dataset will first be examined in its original format and
then reshaped into a structure that is easier to analyze, summarize, and
visualize. Three independent datasets from Discussion 5A will be used in
this project:
Each dataset represents a different data transformation scenario. By restructuring these datasets into tidy formats, the project will demonstrate how appropriate data organization can support meaningful analysis and visualization.
Dataset: COVID World Vaccination Progress
Source: Kaggle
Link: https://www.kaggle.com/datasets/gpreda/covid-world-vaccination-progress
This dataset will be used to examine COVID-19 vaccination progress across countries. It contains several vaccination-related measures, including total vaccinations, the number of people vaccinated, the number of people fully vaccinated, and daily vaccination rates. The dataset will provide an opportunity to demonstrate how multiple vaccination measures can be reorganized into a tidy format. After the transformation, the data will be easier to compare across countries and analyze over time using R.
The dataset contains several vaccination-related measures stored
in separate columns. Although these columns provide useful information,
similar types of measurements are spread across multiple variables
rather than being represented by a single variable that identifies the
vaccination metric. This structure will be treated as a partially wide
format because different vaccination measures are stored in separate
columns. Planned Transformation.
The dataset will be reshaped by combining the vaccination metric columns into a single metric variable and placing their corresponding values into a value variable. This transformation will be performed using the pivot_longer() function from the tidyr package. The resulting tidy dataset will contain variables such as:
• country
• date
• metric
• value
This structure will
make it easier to filter, group, compare, and visualize different
vaccination metrics across countries and over time.
After the dataset is transformed into a tidy format, the analysis will focus on examining vaccination progress across countries and over time. Summary statistics and visualizations will be used to compare vaccination coverage and identify differences between countries. The analysis will include:
• Comparing vaccination coverage across countries
• Examining
changes in vaccination progress over time
• Identifying countries
with relatively higher or lower vaccination coverage
• Creating
visualizations to highlight differences and trends in vaccination
progress
Dataset: Renewable Power Plants / Renewable Capacity Time Series
Source: Kaggle
Link: https://www.kaggle.com/datasets/eugeniyosetrov/renewable-power-plants
This dataset will be used to examine renewable energy generation
across different countries and energy sources. It contains information
related to several types of renewable energy, including solar, wind,
hydro, and other renewable technologies. The dataset provides an
opportunity to examine differences in renewable energy production across
countries and over time.
## Structure Before Tidying
The
dataset contains information organized across multiple columns, with
different renewable energy sources and country-related information
represented separately. Time-related information is also included to
describe changes in energy generation over time. Because related
information is distributed across multiple columns, the dataset will
require restructuring before it can be easily analyzed. The wide format
makes it more difficult to compare energy sources, countries, and time
periods using standard R analysis and visualization functions.
The dataset will be transformed from a wide format into a
longer, tidy format. The transformation will organize the data so that
renewable energy type, country, time period, and generation value are
represented as separate variables.
The transformation will use pivot_longer() from the tidyr
package to combine related columns into appropriate variables. The
resulting tidy dataset will make it easier to group and summarize the
data by country, energy source, and time period.
The transformed dataset will include variables such as:
•
country
• energy_type
• time
• generation
This structure will provide a consistent format for further analysis
and visualization.
After the dataset is transformed into a tidy format, the
analysis will focus on comparing renewable energy generation across
countries, energy sources, and time periods.
The analysis will include:
• Comparing renewable energy
generation across different energy sources
• Comparing renewable
energy output between countries
• Examining changes in renewable
energy generation over time
• Creating visualizations to identify
differences and trends across countries and renewable energy
technologies
These analyses will help demonstrate how tidy data can make it easier to compare renewable energy patterns and identify trends across different countries and energy sources.
Dataset: World GDP by Country (1960–2022)
Source: Kaggle
Link: https://www.kaggle.com/datasets/annafabris/world-gdp-by-country-1960-2022
This dataset will be used to examine GDP trends across countries from
1960 through 2022. It contains GDP values for multiple countries, with
annual values recorded for different years. The dataset will provide an
opportunity to demonstrate how time-based data can be transformed from a
wide format into a tidy structure.
The dataset stores each year as a separate column, creating a
wide-format structure. Instead of having a single variable representing
the year, the year values are used as column names. For example:
Country 1960 1961 1962 1963 … USA value value value value … France value value value value … Japan value value value value …
This structure makes it more difficult to analyze changes over time because the year information is stored in the column names rather than in a separate variable.
The dataset will be transformed from wide format into a tidy long
format by combining the year columns into a single year variable and
placing the corresponding GDP values into a gdp variable.
The
transformation will be performed using the pivot_longer() function from
the tidyr package. The resulting tidy dataset will contain variables
such as:
• country
• year
• gdp
This structure will
make it easier to filter, group, summarize, and visualize GDP data
across countries and years.
After the dataset is transformed into a tidy format, the analysis will focus on examining GDP trends across countries and over time. The analysis will include:
• Comparing GDP trends across selected countries
• Examining
changes in GDP over time
• Calculating and comparing GDP summaries
across different periods
• Identifying countries with notable
changes in GDP
• Creating time-series visualizations to illustrate
economic trends
All data transformations and analyses will be implemented in a Quarto
Markdown document using R and the tidyr, dplyr, and ggplot2 packages.
The workflow for each dataset will include the following steps:
1.
Import the original dataset from a CSV file.
2. Examine the
structure and variables in the raw dataset.
3. Apply pivot_longer()
and other dplyr functions to transform the data into a tidy format.
4. Analyze and summarize the transformed data.
5. Create tables and
visualizations using ggplot2.
The complete workflow will be documented in the Quarto file so that the code and analysis can be reproduced from a clean R session.
The project will demonstrate how wide-format datasets from real-world sources can be transformed into tidy datasets that are easier to analyze and visualize. The transformed datasets will provide a structured format for comparing vaccination progress, examining renewable energy trends, and analyzing economic changes across countries. These transformations will also make it easier to summarize the data, create visualizations, and identify meaningful patterns and trends.
This project will demonstrate how wide-format datasets can be
transformed into tidy datasets using R and the tidyr and dplyr packages.
The three selected datasets COVID-19 vaccination progress, renewable
energy data, and world GDP data have different structures, but each
contains information that can be reorganized to make analysis easier.
Using transformations such as pivot_longer(), the datasets will be
converted into formats where variables are represented as columns and
observations as rows. This will make it easier to filter, summarize,
compare, and visualize the data using dplyr and ggplot2.
Overall, this project will provide practical experience working with
real-world datasets and applying tidy data principles. The final results
will demonstrate how appropriate data transformations can make datasets
easier to understand and prepare them for further analysis and
visualization.