In this assignment, I will use an LLM-generated dataset containing daily sales for three products: Coffee, Tea, and Juice. The dataset includes daily observations from January 1 through March 31, 2026. This dataset provides a time series for multiple items, which will allow me to practice using window functions in R. My plan is to load the dataset into R and use dplyr to organize the data by item and date. I will calculate the year-to-date average, sales for each product and the six-day moving average. I will group the data by item so that the calculations are performed separately for the products.
One challenge I anticipate is making sure that the observations are arranged in the correct date order and grouped by product so that the year-to-date and six-day moving averages are calculated correctly.
Data Source: daily_product_sales_2026.csv
library(dplyr)
## Warning: package 'dplyr' was built under R version 4.4.3
##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
library(slider)
## Warning: package 'slider' was built under R version 4.4.3
I loaded the daily sales dataset into R and examined its structure.
sales_data <- read.csv("daily_product_sales_2026.csv")
str(sales_data)
## 'data.frame': 270 obs. of 3 variables:
## $ date : chr "2026-01-01" "2026-01-02" "2026-01-03" "2026-01-04" ...
## $ item : chr "Coffee" "Coffee" "Coffee" "Coffee" ...
## $ sales: int 118 118 122 131 118 137 119 122 128 129 ...
I converted the date variable to a date format and arranged the observations by product and date. This ensures that the window calculations are performed in chronological order.
sales_data$date <- as.Date(sales_data$date)
sales_data <- sales_data %>%
arrange(item, date)
I grouped the data by product so that the calculations would be performed separately for Coffee, Tea, and Juice. I calculated the year-to-date average which includes all sales observations for a product from the beginning of the year through the current date. I also calculated a six-day moving average, which uses the current day and the previous five days of sales.
sales_data <- sales_data %>%
group_by(item) %>%
mutate(
ytd_average = cummean(sales),
six_day_average = slide_dbl(
sales,
mean,
.before = 5,
.complete = TRUE
)
) %>%
ungroup()
sales_data %>%
group_by(item) %>%
slice_head(n = 10) %>%
ungroup()
## # A tibble: 30 × 5
## date item sales ytd_average six_day_average
## <date> <chr> <int> <dbl> <dbl>
## 1 2026-01-01 Coffee 118 118 NA
## 2 2026-01-02 Coffee 118 118 NA
## 3 2026-01-03 Coffee 122 119. NA
## 4 2026-01-04 Coffee 131 122. NA
## 5 2026-01-05 Coffee 118 121. NA
## 6 2026-01-06 Coffee 137 124 124
## 7 2026-01-07 Coffee 119 123. 124.
## 8 2026-01-08 Coffee 122 123. 125.
## 9 2026-01-09 Coffee 128 124. 126.
## 10 2026-01-10 Coffee 129 124. 126.
## # ℹ 20 more rows
In this analysis, I calculated year-to-date and six-day moving averages for daily sales of Coffee, Tea, and Juice. Grouping the data by product allowed each product’s averages to be calculated independently. The year-to-date average provides a cumulative measure of sales performance, while the six-day moving average focuses on more recent sales.
ChatGPT was used as an AI-assisted tool to generate the artificial daily sales dataset and to assist with code development, debugging, and refinement. Its suggestions and generated code were reviewed, tested, and revised as needed. I remained responsible for the submitted analysis, code, and conclusions.
Tool/model: ChatGPT/ GPT 5.6 Sol Developer: OpenAI Date accessed: September 18, 2026