library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(pastecs)
## 
## Attaching package: 'pastecs'
## 
## The following objects are masked from 'package:dplyr':
## 
##     first, last
## 
## The following object is masked from 'package:tidyr':
## 
##     extract
data <- read_csv("GOVSEMPTIMESERIES.GS00EMP01-2026-09-22T170141.csv")
## Rows: 1976 Columns: 20
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (7): Geographic Area Name (NAME), Meaning of Aggregate Description (AGG_...
## dbl (6): Year (time), Full-Time Employment Coefficient of Variation (FT_EMP_...
## num (7): Full-Time Employment (FT_EMP), Full-Time Payroll (FT_PAY), Part-Tim...
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
data <- data |> rename(TOT_PAY = `Total Full-Time and Part-Time Payroll (TOT_PAY)`)

Question 1 — Select a Variable

Total Full-Time and Part-Time Payroll (TOT_PAY)

Question 2 — Describe the Variable

stat.desc(data$TOT_PAY)
##      nbr.val     nbr.null       nbr.na          min          max        range 
## 1.976000e+03 7.400000e+01 0.000000e+00 0.000000e+00 1.151110e+11 1.151110e+11 
##          sum       median         mean      SE.mean CI.mean.0.95          var 
## 7.139387e+11 2.443218e+07 3.613050e+08 7.303736e+07 1.432384e+08 1.054089e+19 
##      std.dev     coef.var 
## 3.246673e+09 8.985962e+00

Total Full-Time and Part-Time Payroll (TOT_PAY) measures the total amount of payroll paid to full-time and part-time government employees. The variable is measured in dollars and represents the combined payroll for full-time and part-time employees. The dataset contains 1,976 observations for this variable and has no missing values.

Question 3 — Remove NA Values

The variable contains no missing values, so no NA values needed to be removed.

Question 4 — Histogram

hist(data$TOT_PAY)

Because the log transformation cannot be applied to zero, observations with zero payroll were excluded before transforming the variable.

data_clean <- data |> filter(TOT_PAY > 0)

Question 5 — Transform the Variable

data_clean <- data_clean |> mutate(TOT_PAY_LOG = log(TOT_PAY))

Question 6 — Histogram of Transformed Variable

hist(data_clean$TOT_PAY_LOG)