Project 2 — Water Quality

Author

Patricio Romero

Published

October 8, 2026

Introduction

For Project 2, I reviewed the datasets presented in the Week 5 Discussion 5A posts. These posts were used to identify three different datasets that could be prepared and analyzed independently.

This report focuses on the Water Quality dataset. The dataset contains water chemistry measurements that can be used to examine environmental conditions and compare different water quality parameters.

I selected this dataset because water quality is an important environmental topic. The data may help identify differences in chemical measurements across sampling locations.

Dataset Selection

The Water Quality dataset was selected as the second of the three datasets required for Project 2.

Discussion 5A was used to identify the dataset. The original CSV file will be preserved, and the tidying, transformation, and analysis will be completed as part of Project 2.

The objective is to organize the water chemistry measurements into a tidy structure that will make the data easier to compare and analyze.

Data Source

Dataset: Water Quality — Chemistry Data
Original file: Chemistry_Data_Wide_20250715.csv
File format: CSV
Dataset category: Environmental and Water Quality Data

The original CSV file is stored in:

Data/Chemistry_Data_Wide_20250715.csv

The original publisher and official dataset webpage will be added after verifying the source.

Approach

The purpose of this analysis is to compare water chemistry measurements across different sampling locations and identify which parameters show the greatest variation.

The original data will be imported from the CSV file. I will use dplyr to select the required variables, standardize column names, review missing values, and prepare the data for transformation.

The original wide-format dataset will be reviewed to identify the columns containing water chemistry measurements. I will then use pivot_longer() from the tidyr package to transform the measurement columns into a tidy structure.

The transformed dataset will organize the measurements into columns representing the measurement parameter and its recorded value. All analysis and visualizations will use the transformed dataset.

Business Questions

This analysis will address two questions:

  1. How do water chemistry measurements vary across different sampling locations?
  2. Which water chemistry parameters show the greatest variation?

Planned Analysis

The analysis will include:

  • A review of the available water chemistry parameters.
  • A comparison of water chemistry measurements across sampling locations.
  • A calculation of summary statistics for selected chemistry parameters.
  • An examination of which parameters show the greatest variation.
  • Two or three clearly labeled visualizations created with ggplot2.
  • A brief interpretation of the principal findings.

Expected Outcome

The analysis will show how water chemistry measurements differ across the available sampling locations and identify which parameters have the greatest variation.

It will also demonstrate how transforming a wide-format dataset into a tidy structure makes environmental data easier to organize, compare, and analyze.

AI Use

ChatGPT was used to help interpret the assignment requirements, organize the planned approach, improve the English writing, and provide guidance for the R code. I will run, review, and verify the data transformations, analysis, visualizations, and final conclusions.