library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(pastecs)
##
## Attaching package: 'pastecs'
##
## The following objects are masked from 'package:dplyr':
##
## first, last
##
## The following object is masked from 'package:tidyr':
##
## extract
data <- read_csv("GOVSEMPTIMESERIES.GS00EMP01-2026-09-22T170141.csv")
## Rows: 1976 Columns: 20
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (7): Geographic Area Name (NAME), Meaning of Aggregate Description (AGG_...
## dbl (6): Year (time), Full-Time Employment Coefficient of Variation (FT_EMP_...
## num (7): Full-Time Employment (FT_EMP), Full-Time Payroll (FT_PAY), Part-Tim...
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
data <- data |> rename(TOT_PAY = `Total Full-Time and Part-Time Payroll (TOT_PAY)`)
Total Full-Time and Part-Time Payroll (TOT_PAY)
stat.desc(data$TOT_PAY)
## nbr.val nbr.null nbr.na min max range
## 1.976000e+03 7.400000e+01 0.000000e+00 0.000000e+00 1.151110e+11 1.151110e+11
## sum median mean SE.mean CI.mean.0.95 var
## 7.139387e+11 2.443218e+07 3.613050e+08 7.303736e+07 1.432384e+08 1.054089e+19
## std.dev coef.var
## 3.246673e+09 8.985962e+00
Total Full-Time and Part-Time Payroll (TOT_PAY) measures the total amount of payroll paid to full-time and part-time government employees. The variable is measured in dollars and represents the combined payroll for full-time and part-time employees. The dataset contains 1,976 observations for this variable and has no missing values.
The variable contains no missing values, so no NA values needed to be removed.
hist(data$TOT_PAY)
Because the log transformation cannot be applied to zero, observations with zero payroll were excluded before transforming the variable.
data_clean <- data |> filter(TOT_PAY > 0)
data_clean <- data_clean |> mutate(TOT_PAY_LOG = log(TOT_PAY))
hist(data_clean$TOT_PAY_LOG)