DATA 607 Project 2: Data Transformations — Approach

Dataset 1: Student Exam Scores

The initial dataset will have student exam scores through 3 subjects: math, english, and science. Each row will be a student record, with indivudal columns for scores in each subject.

I will firstly format the data into a wide stucture - each subject has its own column. This would create the untidy “issue” where it would be ahrd to compare scores throughout subjects via grouped analysis.

By leveraging tidyr, I will convert the subject columns intro 2 variables: Subject and Score. Dplyr will be used to select important colums, verify vacant values, and categorize observations. Lastly, I will compute the mean scores for every subject using group_by() and summarise().

Dataset 2: Monthly Retail Sales

The 2nd dataset will have monthly sales figures for distinct retail items. Every row will be a record of a product, with seperate columns for monthly sales.

The dataset will firstly be in wide format, with every month indicated by a different column. This format will form the organizational problem for me to solve: I will have to effectively analyze sales trends through months or compare product performances.

I hope to apply pivot_tanger() from tidyr to change the monthly columns into Month and Sales. Then, dplyr will facilitate organization and check vacancies.

Ultimately, I hope to compute total sales for respective items and months usisng group_by() and summarise(). This will let me compare monthly sales trends to pinpoint which products yield the most revenue and if sales go up or down as a function of time.

Dataset 3: NYC Weather

In the midst of the “bipolar weather” at my hometown, I hope to look under the hood, statsitically, at least. I will use weather observations for NYC, including metrics like temp and precipitation values throughout different dates.

The dataset, similar to the other two, will be stored in seperate columns. Based on the source, there can be missing stats or mismatched dating system.

I hope to utilize pivot_longer() from tidyr to restructure measurements intro columns: Measurement_Type and Value. Dplyr will then take important variables, standardize aproppriate dates, filter empty records accordingly, and help total reorganization. In the end, I will compute mean temperatures and overall precipitation through select time intervals .