library(tidyverse)
library(openintro)
data("fastfood", package='openintro')
head(fastfood)
mcdonalds <- fastfood %>%
filter(restaurant == "Mcdonalds")
dairy_queen <- fastfood %>%
filter(restaurant == "Dairy Queen")
Make a plot (or plots) to visualize the distributions of the amount of calories from fat of the options from these two restaurants. How do their centers, shapes, and spreads compare? ### Answer
# Filter data for McDonalds and Dairy Queen
mcdonalds <- fastfood %>% filter(restaurant == "Mcdonalds")
dairy_queen <- fastfood %>% filter(restaurant == "Dairy Queen")
# Combined data for comparative visualization
mc_dq <- fastfood %>% filter(restaurant %in% c("Mcdonalds", "Dairy Queen"))
# Side-by-side / Overlaid Histograms
ggplot(mc_dq, aes(x = cal_fat, fill = restaurant)) +
geom_histogram(binwidth = 50, opacity = 0.6, position = "identity") +
facet_wrap(~restaurant) +
theme_minimal() +
labs(title = "Calories from Fat: McDonald's vs Dairy Queen",
x = "Calories from Fat", y = "Count")
## Warning in geom_histogram(binwidth = 50, opacity = 0.6, position = "identity"):
## Ignoring unknown parameters: `opacity`
# Summary Statistics
mc_dq %>%
group_by(restaurant) %>%
summarise(
mean = mean(cal_fat),
median = median(cal_fat),
sd = sd(cal_fat),
IQR = IQR(cal_fat),
count = n()
)
Center: McDonald’s options tend to have a higher central tendency (mean/median calories from fat) compared to Dairy Queen.
Shape: Both distributions are unimodal and right-skewed. McDonald’s exhibits a stronger right skew with several high-fat extreme values in the upper tail (e.g., large breakfast platters/burgers).
Spread: McDonald’s display a wider spread (larger standard deviation and IQR) than Dairy Queen, indicating higher variability in fat calories across its menu.
dqmean <- mean(dairy_queen$cal_fat)
dqsd <- sd(dairy_queen$cal_fat)
ggplot(data = dairy_queen, aes(x = cal_fat)) +
geom_blank() +
geom_histogram(aes(y = ..density..)) +
stat_function(fun = dnorm, args = c(mean = dqmean, sd = dqsd), col = "tomato")
## Warning: The dot-dot notation (`..density..`) was deprecated in ggplot2 3.4.0.
## ℹ Please use `after_stat(density)` instead.
## This warning is displayed once per session.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.
## `stat_bin()` using `bins = 30`. Pick better value `binwidth`.
Based on the this plot, does it appear that the data follow a nearly normal distribution? ### Answer No, not perfectly, but it is nearly normal. While the density histogram roughly forms a unimodal bell shape matching the overlaid red normal curve around the peak, there is noticeable right-skewness. The distribution has a longer right tail and slight discrepancies around the center-peak height compared to the ideal theoretical normal curve.
Make a normal probability plot of sim_norm. Do all of the points fall on the line? How does this plot compare to the probability plot for the real data? (Since sim_norm is not a data frame, it can be put directly into the sample argument and the data argument can be dropped.) ### Answer
Do all points fall on the line? No. Even data drawn from a true normal distribution exhibits random sampling variability, causing minor wiggles or deviations—especially at the extreme upper and lower tails.
Comparison: The simulated Q-Q plot follows the diagonal line much tighter overall than the real Dairy Queen data, which shows a pronounced upward curve at the upper tail (characteristic of right skewness).
dqmean <- mean(dairy_queen$cal_fat)
dqsd <- sd(dairy_queen$cal_fat)
# Generate simulated normal data
sim_norm <- rnorm(n = nrow(dairy_queen), mean = dqmean, sd = dqsd)
# Normal Q-Q plot for simulated data
ggplot(data = NULL, aes(sample = sim_norm)) +
geom_line(stat = "qq") +
labs(title = "Normal Q-Q Plot of Simulated Data",
x = "Theoretical Quantiles", y = "Sample Quantiles")
qqnormsim(sample = cal_fat, data = dairy_queen)
## Warning: `aes_string()` was deprecated in ggplot2 3.0.0.
## ℹ Please use tidy evaluation idioms with `aes()`.
## ℹ See also `vignette("ggplot2-in-packages")` for more information.
## ℹ The deprecated feature was likely used in the openintro package.
## Please report the issue at
## <https://github.com/OpenIntroStat/openintro/issues>.
## This warning is displayed once per session.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.
Does the normal probability plot for the calories from fat look similar to the plots created for the simulated data? That is, do the plots provide evidence that the calories are nearly normal? ### Answer Yes, it appears nearly normal. While the real Dairy Queen data shows some slight deviation (points bending upward in the upper tail), its pattern is reasonably consistent with the random variation observed across the 8 simulated plots generated by qqnormsim(). Because real-world sample data naturally fluctuates, the level of curvature in the actual data is small enough to treat the distribution as approximately normal.
Using the same technique, determine whether or not the calories from McDonald’s menu appear to come from a normal distribution. ### Answer No, McDonald’s fat calories do not follow a normal distribution. The data is strongly right-skewed. The real Q-Q plot exhibits a distinct convex shape (points sharply curve upward on the right side above the line), which is noticeably more extreme than any of the simulated normal plots.
# Density Histogram overlay
mc_mean <- mean(mcdonalds$cal_fat)
mc_sd <- sd(mcdonalds$cal_fat)
ggplot(mcdonalds, aes(x = cal_fat)) +
geom_histogram(aes(y = ..density..), binwidth = 50) +
stat_function(fun = dnorm, args = list(mean = mc_mean, sd = mc_sd), col = "red")
# Q-Q plot simulation
qqnormsim(sample = cal_fat, data = mcdonalds)
Write out two probability questions that you would like to answer about any of the restaurants in this dataset. Calculate those probabilities using both the theoretical normal distribution as well as the empirical distribution (four probabilities in all). Which one had a closer agreement between the two methods? ### Answer Question A: Probability a McDonald’s menu item has more than 500 calories. Question B: Probability a Dairy Queen menu item has less than 200 calories from fat.
Dairy Queen’s fat calories (< 200) yields a closer agreement between theoretical and empirical calculations. Because Dairy Queen’s fat calorie distribution is much closer to a normal curve than McDonald’s highly skewed total calorie count, its theoretical probability matches the empirical percentage far more accurately.
# --- Question A: McDonald's total calories > 500 ---
mc_cal_mean <- mean(mcdonalds$calories)
mc_cal_sd <- sd(mcdonalds$calories)
# Theoretical
mc_theo <- 1 - pnorm(500, mean = mc_cal_mean, sd = mc_cal_sd)
# Empirical
mc_emp <- mcdonalds %>%
summarise(percent = mean(calories > 500)) %>%
pull(percent)
# --- Question B: Dairy Queen calories from fat < 200 ---
# Theoretical
dq_theo <- pnorm(200, mean = dqmean, sd = dqsd)
# Empirical
dq_emp <- dairy_queen %>%
summarise(percent = mean(cal_fat < 200)) %>%
pull(percent)
# Summary of results
tibble(
Question = c("McDonald's Cal > 500", "Dairy Queen Fat Cal < 200"),
Theoretical = c(mc_theo, dq_theo),
Empirical = c(mc_emp, dq_emp),
Difference = abs(c(mc_theo, dq_theo) - c(mc_emp, dq_emp))
)
Now let’s consider some of the other variables in the dataset. Out of all the different restaurants, which ones’ distribution is the closest to normal for sodium? ### Answer Arby’s (or Burger King, depending on exact binning/sample size interpretation) exhibits the Q-Q plot where sodium data points stick closest to a straight line. Chains like Chick-fil-A or Subway feature heavier tail deviations or discrete clustering.
# Visualize Q-Q plots for sodium across all restaurants
ggplot(fastfood, aes(sample = sodium)) +
geom_line(stat = "qq") +
facet_wrap(~restaurant, scales = "free") +
theme_minimal() +
labs(title = "Sodium Normal Q-Q Plots by Restaurant")
Note that some of the normal probability plots for sodium distributions seem to have a stepwise pattern. why do you think this might be the case? ### Answer The stepwise (stair-case) pattern occurs because sodium values are often rounded (e.g., reported to the nearest 10 mg or 50 mg on nutritional panels) or because multiple menu items share identical standardized sodium levels (e.g., identical condiments or base ingredients). Discrete or heavily rounded data creates repeated ties in values, producing horizontal plateaus on a Q-Q plot.
As you can see, normal probability plots can be used both to assess normality and visualize skewness. Make a normal probability plot for the total carbohydrates from a restaurant of your choice. Based on this normal probability plot, is this variable left skewed, symmetric, or right skewed? Use a histogram to confirm your findings. ### Answer Q-Q Plot Interpretation: The points form an upward curving concave arc (bending upward at the right extreme above the line), indicating that the sample quantiles increase much faster than theoretical normal quantiles at the upper end. This signifies that the variable is right-skewed.
Histogram Confirmation: The histogram confirms this finding—it displays a prominent peak on the left side (around 30–50g) with a long tail extending toward higher carbohydrate counts (80g+).
# Choice of restaurant: Arby's
arbys <- fastfood %>% filter(restaurant == "Arbys")
# 1. Normal Q-Q Plot
ggplot(arbys, aes(sample = total_carb)) +
geom_line(stat = "qq") +
labs(title = "Q-Q Plot: Arby's Total Carbs",
x = "Theoretical Quantiles", y = "Sample Quantiles")
# 2. Histogram Confirmation
ggplot(arbys, aes(x = total_carb)) +
geom_histogram(binwidth = 10, fill = "steelblue", color = "white") +
labs(title = "Histogram: Arby's Total Carbs",
x = "Total Carbohydrates (g)", y = "Count")