library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(pastecs)
##
## Attaching package: 'pastecs'
##
## The following objects are masked from 'package:dplyr':
##
## first, last
##
## The following object is masked from 'package:tidyr':
##
## extract
For this analysis, I am interested in the number of victims reported in completed child abuse and neglect investigations.
The X..of.Victims variable measures the number of victims reported in each county, fiscal year, and program observation. The data include observations from Texas counties across different fiscal years and programs.
dfps_data<-read.csv("dfps_data.csv")
head(dfps_data)
## Fiscal.Year Program County Region Confirmed.Victims X..of.Victims
## 1 2025 DCI Hutchinson 1-Lubbock Not Confirmed 1
## 2 2025 DCI Lubbock 1-Lubbock Not Confirmed 36
## 3 2025 DCI Lubbock 1-Lubbock Confirmed Victim 4
## 4 2025 DCI Ochiltree 1-Lubbock Not Confirmed 1
## 5 2025 DCI Potter 1-Lubbock Not Confirmed 13
## 6 2025 DCI Potter 1-Lubbock Confirmed Victim 1
str(dfps_data)
## 'data.frame': 3121 obs. of 6 variables:
## $ Fiscal.Year : int 2025 2025 2025 2025 2025 2025 2025 2025 2025 2025 ...
## $ Program : chr "DCI" "DCI" "DCI" "DCI" ...
## $ County : chr "Hutchinson" "Lubbock" "Lubbock" "Ochiltree" ...
## $ Region : chr "1-Lubbock" "1-Lubbock" "1-Lubbock" "1-Lubbock" ...
## $ Confirmed.Victims: chr "Not Confirmed" "Not Confirmed" "Confirmed Victim" "Not Confirmed" ...
## $ X..of.Victims : chr "1" "36" "4" "1" ...
The X..of.Victims variable is stored as character data, so it must be converted to numeric data before statistical calculations can be performed.
dfps_data<-dfps_data %>% mutate(X..of.Victims=as.numeric(X..of.Victims))
## Warning: There was 1 warning in `mutate()`.
## ℹ In argument: `X..of.Victims = as.numeric(X..of.Victims)`.
## Caused by warning:
## ! NAs introduced by coercion
stat.desc(dfps_data$X..of.Victims,norm=T)
## nbr.val nbr.null nbr.na min max range
## 3.116000e+03 0.000000e+00 5.000000e+00 1.000000e+00 9.210000e+02 9.200000e+02
## sum median mean SE.mean CI.mean.0.95 var
## 6.072800e+04 4.000000e+00 1.948909e+01 9.373862e-01 1.837957e+00 2.738007e+03
## std.dev coef.var skewness skew.2SE kurtosis kurt.2SE
## 5.232597e+01 2.684885e+00 6.826962e+00 7.782686e+01 6.772478e+01 3.861523e+02
## normtest.W normtest.p
## 3.596936e-01 3.988997e-74
The variable contains 3,116 observations after missing values are removed. The mean number of victims is 19.49, while the median is 4. The values range from 1 to 921. The mean is substantially higher than the median, indicating that the distribution is strongly right-skewed.
sum(is.na(dfps_data$X..of.Victims))
## [1] 5
There were 5 missing values in the number of victims variable.
dfps_data_cleaned<-dfps_data %>% filter(!is.na(X..of.Victims))
sum(is.na(dfps_data_cleaned$X..of.Victims))
## [1] 0
hist(dfps_data_cleaned$X..of.Victims,breaks=10,probability = T)
lines(density(dfps_data_cleaned$X..of.Victims),col='red',lwd=2)
The histogram shows that the number of victims is strongly right-skewed. Most observations have relatively small numbers of victims, while a smaller number of observations have substantially larger victim counts.
Because the number of victims variable is strongly right-skewed, a log transformation was used. The log transformation can reduce the influence of larger observations and reduce skew in the distribution.
dfps_data_log<-dfps_data_cleaned %>% mutate(LOG_VICTIMS=log(X..of.Victims)) %>% select(X..of.Victims,LOG_VICTIMS)
head(dfps_data_log)
## X..of.Victims LOG_VICTIMS
## 1 1 0.000000
## 2 36 3.583519
## 3 4 1.386294
## 4 1 0.000000
## 5 13 2.564949
## 6 1 0.000000
hist(log(dfps_data_cleaned$X..of.Victims),breaks=10,probability = T)
lines(density(log(dfps_data_cleaned$X..of.Victims)),col='red',lwd=2)
The log transformation reduced the influence of the larger victim counts and made the distribution less right-skewed.