Loading the Data

setwd("~/Documents/R Files/Homework 4")
districts <- read_csv("district_info.csv", skip=1)

1) Select the Variable #For this analysis, I selected total student enrollment DPETALLC from the 2024–25 Texas Academic Performance Reports district-level dataset. This variable measures the total number of students enrolled in each Texas public school district. I selected enrollment because district sizes vary considerably across Texas, ranging from very small rural districts to very large urban districts.

2) Use pastecs::stat.desc to describe the variable. Include a few sentences about what the variable is and what it’s measuring.

stat.desc(districts$DPETALLC)
     nbr.val     nbr.null       nbr.na          min          max        range 
1.208000e+03 0.000000e+00 0.000000e+00 9.000000e+00 1.760390e+05 1.760300e+05 
         sum       median         mean      SE.mean CI.mean.0.95          var 
5.530499e+06 8.870000e+02 4.578228e+03 3.552110e+02 6.968997e+02 1.524192e+08 
     std.dev     coef.var 
1.234582e+04 2.696637e+00 

#The variable DPETALLC measures total student enrollment for each Texas public school district. There are 1,208 districts in the dataset, with enrollment ranging from 9 students to 176,039 students. The median district enrollment is 887 students, while the mean is approximately 4,578 students. The large difference between the mean and median suggests that the distribution is right-skewed, with a small number of very large districts increasing the overall average.

3) Remove NA’s if needed using dplyr:filter (or anything similar)

districts_clean <- filter(districts, !is.na(DPETALLC))
cat(nrow(districts_clean))
1208

#As shown in step 2, there were no missing observations for total student enrollment, so no observations need to be removed.

4) Provide a histogram of the variable (as shown in the lesson)

hist(districts$DPETALLC,
     main = "Distribution of Texas School District Enrollment",
     xlab = "Total Student Enrollment",
     col = "lightsteelblue",
     border = "gray30",
     xaxt = "n")

#I used the function below to add commas to the numbers where necessary.
axis(1,
     at = axTicks(1),
     labels = format(axTicks(1), big.mark = ",", scientific = FALSE))

5) Transform the variable using the log transformation or square root transformation (whatever is more appropriate) using dplyr::mutate or something similar

districts_clean <- mutate(
  districts_clean,
  log_enrollment = log(DPETALLC)
)

6) provide a histogram of the transformed variable

hist(districts_clean$log_enrollment,
     main = "Distribution of Log-Transformed District Enrollment",
     xlab = "Log of Total Student Enrollment",
     col = "lightsteelblue",
     border = "gray30")