Include the R commands/functions that you used to find your answer (show R chunk). Answers without supporting code will not receive credit. Write full sentences to describe your findings.
# Load the package
library(tidyverse)
Dataset: Remind us which dataset(s) you are exploring. Download the dataset from the City of Austin Data Portal as a .csv file and import it into R. Report the number of rows and columns in the dataset. Show the first few rows and the main variables of interest. If you are using more than one dataset, explain how you would join the datasets.
# Import the dataset and name is as permits
permits <- read_csv("Issued_Construction_Permits_20261007(in).csv", na="")
## Warning: One or more parsing issues, call `problems()` on your data frame for details,
## e.g.:
## dat <- vroom(...)
## problems(dat)
## Rows: 4774 Columns: 8
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (5): Condominium, Issuance.Method, Status.Current, Completed.Date, Issue...
## dbl (2): Number.Of.Floors, Housing.Units
## num (1): Total.New.Add.SQFT
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
# Taking a look inside the data
head(permits)
## # A tibble: 6 × 8
## Condominium Issuance.Method Status.Current Completed.Date Number.Of.Floors
## <chr> <chr> <chr> <chr> <dbl>
## 1 No Permit Center Final 8/19/2026 1
## 2 No Permit Center Final 9/28/2026 1
## 3 No Permit Center Final 6/2/2026 1
## 4 No Permit Center Final 9/4/2026 1
## 5 No Permit Center Final 9/25/2026 1
## 6 No Permit Center Final 10/2/2026 1
## # ℹ 3 more variables: Housing.Units <dbl>, Issued.Date <chr>,
## # Total.New.Add.SQFT <dbl>
Research Question: Remind us of your research question(s).
Research Question: How does job size and type impact how long it takes to complete a permit, and how has this changed over time?
Variables: Identify the variables that you plan to use to answer your research question. For each variable:
Briefly describe what it represents and whether you will treat it as numeric or categorical.
Examine its distribution using appropriate visualizations and summary statistics.
As you explore your variables, look for anything that may need to be addressed before your analysis. For example: Are numeric variables stored as numbers? Are there categories that should be combined or renamed? Are there missing or unusual values? Would creating a new variable help answer your research question?
Condominium:
It represent if it’s a condo (yes or no)
We are treating this as a categorical nominal variable
# The distribution of Condominium
permits |>
filter(!is.na(Condominium)) |>
ggplot() +
geom_bar(aes(x = Condominium)) +
labs(
title = "Total Count of Condominum",
x = "Condominium",
y = ""
)
# Summary statistic for Condominium
permits |>
filter(!is.na(Condominium)) |>
group_by(Condominium) |>
summarize(
count = n()
)
## # A tibble: 2 × 2
## Condominium count
## <chr> <int>
## 1 No 4712
## 2 Yes 62
Condominium is a categorical variable, so we summarize it with counts: 4712 No and 62 Yes.
Total new add SQFT
Additional square footage gained from work being done
Numeric continuous
# ggplot for Total New Add SQFT
permits |>
ggplot(aes(x = Total.New.Add.SQFT)) +
geom_histogram() +
# Scale the x axis logarithmically
scale_x_log10() +
# Add labels
labs(
title = "Distribution of Total Newly Added Square Footage",
x = "Total New Add (sq. ft)",
y = "Number of permits"
)
## Warning in scale_x_log10(): log-10 transformation introduced infinite values.
## `stat_bin()` using `bins = 30`. Pick better value `binwidth`.
## Warning: Removed 499 rows containing non-finite outside the scale range
## (`stat_bin()`).
# Summary statistic for Total New Add SQFT
permits |>
summarize(
number_observed = n(),
mean_sqft = mean(Total.New.Add.SQFT),
sd_sqft = sd(Total.New.Add.SQFT),
median_sqft = median(Total.New.Add.SQFT),
IQR_sqft = IQR(Total.New.Add.SQFT)
)
## # A tibble: 1 × 5
## number_observed mean_sqft sd_sqft median_sqft IQR_sqft
## <int> <dbl> <dbl> <dbl> <dbl>
## 1 4774 2442. 10774. 1816. 2450
Issuance Method:
“Permit Center” represents any permit which a permit center technician manually issued
“Online” includes permits issued through the website including Austin Build and Connect or Customer Self Assignment
Categorical
# ggplot for Issuance Method
permits |>
group_by(Issuance.Method) |>
ggplot(aes(x = Issuance.Method)) +
# Add
geom_bar() +
# Add labels
labs(
x = "Issuance Method",
y = "Count",
title = "Counts of Issuance Methods"
)
# Summary statistic for Issuance Method
permits |>
group_by(Issuance.Method) |>
summarize(
number_observed = n()
)
## # A tibble: 1 × 2
## Issuance.Method number_observed
## <chr> <int>
## 1 Permit Center 4774
Issuance Method is a categorical variable, so we summarize it with counts: 4774 Permit Center.
Status Current:
Current status of permit
Categorical nominal
# ggplot for Status Current
permits |>
group_by(Status.Current) |>
ggplot(aes(x = Status.Current)) +
# Add
geom_bar() +
# Add labels
labs(
x = "Current Status",
y = "Count",
title = "Counts of Current Statuses"
)
# Summary Statistic for Status Current
permits |>
group_by(Status.Current) |>
summarize(
number_observed = n()
)
## # A tibble: 1 × 2
## Status.Current number_observed
## <chr> <int>
## 1 Final 4774
Current Status is a categorical variable, so we summarize it with counts: 4774 Final.
Completed Date:
Date on which the permit was completed
Categorical nominal variable
# Create a histogram for Completed Date
Number of Floors:
How many floors the property has
Discrete numeric variable
# Create a histogram for Number of Floors
The distribution of Number of Floors is right-skewed, so the appropriate statistics are the median and IQR. The median is 1 floor and the IQR is 1 floor.
Housing Units:
Number of household units for a given building
Discrete numeric variable
# Create a histogram for Housing Units
ggplot(data = permits) +
geom_histogram(
aes(x = Housing.Units),
binwidth = 1, center = 0.5,
color = "black", fill = "lightblue"
) +
scale_x_continuous(limits = c(0,10), breaks = seq(0, 10, 1)) +
labs(
title = "Distribution of Housing Units",
x = "Housing Units",
y = "count"
)
## Warning: Removed 14 rows containing non-finite outside the scale range
## (`stat_bin()`).
# Summary statistc for Housing Units
permits |>
mutate(Housing.Units = as.numeric(Housing.Units)) |>
summarize(
number_observed = n(),
mean_housing_unit = mean(Housing.Units, na.rm = TRUE),
sd_housing_unit = sd(Housing.Units, na.rm = TRUE),
median_housing_unit = median(Housing.Units, na.rm = TRUE),
IQR_housing_unit = IQR(Housing.Units, na.rm = TRUE)
)
## # A tibble: 1 × 5
## number_observed mean_housing_unit sd_housing_unit median_housing_unit
## <int> <dbl> <dbl> <dbl>
## 1 4774 1.07 3.03 1
## # ℹ 1 more variable: IQR_housing_unit <dbl>
The distribution of Housing Units is right-skewed, so the appropriate statistics are the median and IQR. The median is 1 unit and the IQR is 0 unit.
Issued Date:
Date on which the permit was issued
Categorical nominal variable
# Create a histogram for Issued Date
Some of the variables are dates, so they need to be handled differently than numbers. Creating a variable that stores the time between issuance and completion would help us answer the research question. For missing values, we can try out different policies of handling missing values (like dropping them or not) and see which one performs the best. We also noticed that there were no permits with Issuance Method of Online in the year 2026.
Wrangling: Based on what you discovered about your variables, prepare your dataset for analysis by performing any data wrangling that is appropriate for your project. This might include:
Then briefly summarize:
We filtered for where the total new add square feet is not null, the current status is “Final”, and the Calendar Year Issued is 2026. We also added a variable called Time.Taken which stores the number of days between the issued date and completed date
We will use the variables discussed above.
After filtering, there were no permits where the Issuance Method was “Online”. This may be an issue later on which we will need to consider. There may also be more data wrangling we need to do that we haven’t done yet.