library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(readxl)
library(pastecs)
##
## Attaching package: 'pastecs'
##
## The following objects are masked from 'package:dplyr':
##
## first, last
##
## The following object is masked from 'package:tidyr':
##
## extract
library(car)
## Loading required package: carData
##
## Attaching package: 'car'
##
## The following object is masked from 'package:dplyr':
##
## recode
##
## The following object is masked from 'package:purrr':
##
## some
Cities<-read_xlsx("FiSC-Full-Dataset-2023-Update.xlsx")
stat.desc(Cities$education_city)
## nbr.val nbr.null nbr.na min max range
## 1.024600e+04 3.180000e+03 0.000000e+00 0.000000e+00 6.196860e+03 6.196860e+03
## sum median mean SE.mean CI.mean.0.95 var
## 4.392838e+06 2.968000e+01 4.287369e+02 9.666649e+00 1.894852e+01 9.574283e+05
## std.dev coef.var
## 9.784827e+02 2.282245e+00
The variable I have chosen for this excercise is the spending for all education of cities within the data set. This is to include K-12 education, higher education, and libraries. The values represented are per capita dollars, so the value given is spending per resident of each city. The range of the values is fairly large, with the maximum spending recorded 6196.86 per capita, and the lowest being 0. It would seem that most values tend towartds the lower end, as the median is 29.68. Even so, it would seem that the vlaues are right skewed, as the mean is pulled higher to 428.73.
The following histogram shows the frequency of amounts of spending, with the zeros included
hist(Cities$education_city)
The histogram shows that an overwhelming majority of cities in this dataset do not supply any additional spending towards education.Miuch of this can be explained by school districts providing the majority of funding, which is also included in the data set. So, we can remove the zeros in the data to show only the municipalities that send additional funds on education.
Cities_clean<-Cities %>% drop_na(education_city) %>% filter(education_city>0)
hist(Cities_clean$education_city)
Even with all the zeros taken out from the data, the histogram has not changed significantly, showing the majority of values are still on the very low end.
Cities_clean<-Cities_clean %>% mutate(education_city_SQRT=sqrt(education_city))
hist(Cities_clean$education_city_SQRT)
Squaring the values does some to help differentiate the degrees of municiple education spending, but not by a large degree. It is still largely right scewed, but there is some difference in the lower end values. It seems that while a great many cities choose to spend very little on education, a larger porportion choose to spend slightly more, though is still dwarfed by a few number of cities.