library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(pastecs)
## 
## Attaching package: 'pastecs'
## 
## The following objects are masked from 'package:dplyr':
## 
##     first, last
## 
## The following object is masked from 'package:tidyr':
## 
##     extract

Number of Victims

For this analysis, I am interested in the number of victims reported in completed child abuse and neglect investigations.

The X..of.Victims variable measures the number of victims reported in each county, fiscal year, and program observation. The data include observations from Texas counties across different fiscal years and programs.

dfps_data<-read.csv("dfps_data.csv")
head(dfps_data)
##   Fiscal.Year Program     County    Region Confirmed.Victims X..of.Victims
## 1        2025     DCI Hutchinson 1-Lubbock     Not Confirmed             1
## 2        2025     DCI    Lubbock 1-Lubbock     Not Confirmed            36
## 3        2025     DCI    Lubbock 1-Lubbock  Confirmed Victim             4
## 4        2025     DCI  Ochiltree 1-Lubbock     Not Confirmed             1
## 5        2025     DCI     Potter 1-Lubbock     Not Confirmed            13
## 6        2025     DCI     Potter 1-Lubbock  Confirmed Victim             1
str(dfps_data)
## 'data.frame':    3121 obs. of  6 variables:
##  $ Fiscal.Year      : int  2025 2025 2025 2025 2025 2025 2025 2025 2025 2025 ...
##  $ Program          : chr  "DCI" "DCI" "DCI" "DCI" ...
##  $ County           : chr  "Hutchinson" "Lubbock" "Lubbock" "Ochiltree" ...
##  $ Region           : chr  "1-Lubbock" "1-Lubbock" "1-Lubbock" "1-Lubbock" ...
##  $ Confirmed.Victims: chr  "Not Confirmed" "Not Confirmed" "Confirmed Victim" "Not Confirmed" ...
##  $ X..of.Victims    : chr  "1" "36" "4" "1" ...

The X..of.Victims variable is stored as character data, so it must be converted to numeric data before statistical calculations can be performed.

dfps_data<-dfps_data %>% mutate(X..of.Victims=as.numeric(X..of.Victims))
## Warning: There was 1 warning in `mutate()`.
## ℹ In argument: `X..of.Victims = as.numeric(X..of.Victims)`.
## Caused by warning:
## ! NAs introduced by coercion

Descriptive Statistics

stat.desc(dfps_data$X..of.Victims,norm=T)
##      nbr.val     nbr.null       nbr.na          min          max        range 
## 3.116000e+03 0.000000e+00 5.000000e+00 1.000000e+00 9.210000e+02 9.200000e+02 
##          sum       median         mean      SE.mean CI.mean.0.95          var 
## 6.072800e+04 4.000000e+00 1.948909e+01 9.373862e-01 1.837957e+00 2.738007e+03 
##      std.dev     coef.var     skewness     skew.2SE     kurtosis     kurt.2SE 
## 5.232597e+01 2.684885e+00 6.826962e+00 7.782686e+01 6.772478e+01 3.861523e+02 
##   normtest.W   normtest.p 
## 3.596936e-01 3.988997e-74

The variable contains 3,116 observations after missing values are removed. The mean number of victims is 19.49, while the median is 4. The values range from 1 to 921. The mean is substantially higher than the median, indicating that the distribution is strongly right-skewed.

Missing Values

sum(is.na(dfps_data$X..of.Victims))
## [1] 5

There were 5 missing values in the number of victims variable.

dfps_data_cleaned<-dfps_data %>% filter(!is.na(X..of.Victims))
sum(is.na(dfps_data_cleaned$X..of.Victims))
## [1] 0

Histogram

hist(dfps_data_cleaned$X..of.Victims,breaks=10,probability = T)
lines(density(dfps_data_cleaned$X..of.Victims),col='red',lwd=2)

The histogram shows that the number of victims is strongly right-skewed. Most observations have relatively small numbers of victims, while a smaller number of observations have substantially larger victim counts.

Transformation

Because the number of victims variable is strongly right-skewed, a log transformation was used. The log transformation can reduce the influence of larger observations and reduce skew in the distribution.

dfps_data_log<-dfps_data_cleaned %>% mutate(LOG_VICTIMS=log(X..of.Victims)) %>% select(X..of.Victims,LOG_VICTIMS)

Number of Victims After Log Transformation

head(dfps_data_log)
##   X..of.Victims LOG_VICTIMS
## 1             1    0.000000
## 2            36    3.583519
## 3             4    1.386294
## 4             1    0.000000
## 5            13    2.564949
## 6             1    0.000000

Histogram of Transformed Variable

hist(log(dfps_data_cleaned$X..of.Victims),breaks=10,probability = T)
lines(density(log(dfps_data_cleaned$X..of.Victims)),col='red',lwd=2)

The log transformation reduced the influence of the larger victim counts and made the distribution less right-skewed.