Include the R commands/functions that you used to find your answer (show R chunk). Answers without supporting code will not receive credit. Write full sentences to describe your findings.
# Load the package
library(tidyverse)
Dataset: Remind us which dataset(s) you are exploring. Download the dataset from the City of Austin Data Portal as a .csv file and import it into R. Report the number of rows and columns in the dataset. Show the first few rows and the main variables of interest. If you are using more than one dataset, explain how you would join the datasets.
permits <- read.csv("Issued_Construction_Permits_20261007.csv", na="")
head(permits)
## Condominium Issuance.Method Status.Current Completed.Date Number.Of.Floors
## 1 No Permit Center Final 2026/08/19 1
## 2 No Permit Center Final 2026/09/28 1
## 3 No Permit Center Final 2026/06/02 1
## 4 No Permit Center Final 2026/09/04 1
## 5 No Permit Center Final 2026/09/25 1
## 6 No Permit Center Final 2026/10/02 1
## Housing.Units Issued.Date Total.New.Add.SQFT
## 1 1 2026/04/08 2,062
## 2 1 2026/03/06 5,000
## 3 1 2026/02/25 2,446
## 4 1 2026/01/23 1,800
## 5 1 2026/09/02 114
## 6 2 2026/09/26 1,500
Research Question: Remind us of your research question(s).
Research Question: How does job size and type impact how long it takes to complete a permit, and how has this changed over time?
Variables: Identify the variables that you plan to use to answer your research question. For each variable:
Briefly describe what it represents and whether you will treat it as numeric or categorical.
Examine its distribution using appropriate visualizations and summary statistics.
As you explore your variables, look for anything that may need to be addressed before your analysis. For example: Are numeric variables stored as numbers? Are there categories that should be combined or renamed? Are there missing or unusual values? Would creating a new variable help answer your research question?
Condominium:
It represent if it’s a condo (yes or no)
We are treating this as a categorical nominal variable
# The distribution of Condominium
permits |>
filter(!is.na(Condominium)) |>
ggplot() +
geom_bar(aes(x = Condominium)) +
labs(
x = "Condominium",
y = ""
)
# Summarize Condominium
permits |>
filter(!is.na(Condominium)) |>
group_by(Condominium) |>
summarize(
count = n()
)
## # A tibble: 2 × 2
## Condominium count
## <chr> <int>
## 1 No 4712
## 2 Yes 62
Condominium is a categorical variable, so we summarize it using counts. There are far fewer permits for condominiums than for other types of permits.
Total new add SQFT
Additional square footage gained from work being done
Numeric continuous
permits |>
mutate(Total.New.Add.SQFT = parse_number(Total.New.Add.SQFT)) |>
ggplot(aes(x = Total.New.Add.SQFT)) +
geom_histogram() +
# Scale the x axis logarithmically
scale_x_log10() +
# Add labels
labs(
title = "Distribution of Total newly added square footage",
x = "Total New Add (sq. ft)",
y = "Number of permits"
)
## Warning in scale_x_log10(): log-10 transformation introduced infinite values.
## `stat_bin()` using `bins = 30`. Pick better value `binwidth`.
## Warning: Removed 499 rows containing non-finite outside the scale range
## (`stat_bin()`).
# Summary statistics for Total New Add SQFT
permits |>
mutate(Total.New.Add.SQFT = parse_number(Total.New.Add.SQFT)) |>
summarize(
number_observed = n(),
mean_sqft = mean(Total.New.Add.SQFT),
sd_sqft = sd(Total.New.Add.SQFT),
median_sqft = median(Total.New.Add.SQFT),
IQR_sqft = IQR(Total.New.Add.SQFT)
)
## number_observed mean_sqft sd_sqft median_sqft IQR_sqft
## 1 4774 2442.205 10774.18 1815.5 2450
The total new add sq. ft. seems to be centered around 2000, with a mean of 2442 (and SD 10774.18) and a median of 1815 (with an IQR of 2450). This data is spread out a lot as shown by the SD and and IQR.
Issuance Method:
“Permit Center” represents any permit which a permit center technician manually issued
“Online” includes permits issued through the website including Austin Build and Connect or Customer Self Assignment
Categorical
permits |>
group_by(Issuance.Method) |>
ggplot(aes(x = Issuance.Method)) +
# Add bar plot
geom_bar() +
# Add labels
labs(
x = "Issuance Method",
y = "Count",
title = "Counts of Issuance Methods"
)
# Summary for Issuance Method
permits |>
group_by(Issuance.Method) |>
summarize(
number_observed = n()
)
## # A tibble: 1 × 2
## Issuance.Method number_observed
## <chr> <int>
## 1 Permit Center 4774
Issuance Method is a categorical variable, so we summarize it with counts. Interestingly, there are no records with Issuance Method of “Online” in the dataset.
Status Current:
Current status of permit
Categorical nominal
permits |>
group_by(Status.Current) |>
ggplot(aes(x = Status.Current)) +
# Add bar plot
geom_bar() +
# Add labels
labs(
x = "Current Status",
y = "Count",
title = "Counts of Current Statuses"
)
# Summary for Current Status
permits |>
group_by(Status.Current) |>
summarize(
number_observed = n()
)
## # A tibble: 1 × 2
## Status.Current number_observed
## <chr> <int>
## 1 Final 4774
Current status is a categorical variable, so we summarize it with counts. The dataset is pre-filtered to only include records whose status is “Final”.
Completed Date:
Date on which the permit was completed
Discrete numeric variable (months are categorical ordinal)
# Create a histogram for Completed Date (by month)
permits |>
mutate(month_label = month(Completed.Date, label = TRUE)) |>
ggplot() +
geom_bar(aes(x = month_label)) +
labs(
x = "Month",
y = "Count",
title = "Permits completed per month"
)
# Summary for Date Completed (by month)
permits |>
mutate(month_label = month(Completed.Date, label = TRUE)) |>
group_by(month_label) |>
summarize(
num_observed = n()
)
## # A tibble: 12 × 2
## month_label num_observed
## <ord> <int>
## 1 Jan 34
## 2 Feb 99
## 3 Mar 257
## 4 Apr 317
## 5 May 485
## 6 Jun 728
## 7 Jul 800
## 8 Aug 864
## 9 Sep 1009
## 10 Oct 176
## 11 Nov 1
## 12 Dec 4
The month of the completed date is a categorical ordinal variable, so we summarize it using counts.
Number of Floors:
How many floors the property has
Discrete numeric variable
# Create a histogram for Number of Floors
permits |>
mutate(Number.Of.Floors = parse_number(Number.Of.Floors)) |>
ggplot() +
# Add a histogram
geom_histogram(
aes(x = Number.Of.Floors),
binwidth = 1, center = 0.5,
color = "black", fill = "lightblue"
) +
# Scale the graph from 0 to 10
scale_x_continuous(limits = c(0,10), breaks = seq(0, 10, 1)) +
# Add labels
labs(
title = "Distribution of Number of Floors",
x = "Number of Floors",
y = "count"
)
## Warning: Removed 14 rows containing non-finite outside the scale range
## (`stat_bin()`).
# Summary for number of floors
permits |>
mutate(Number.Of.Floors = parse_number(Number.Of.Floors)) |>
summarize(
number_observed = n(),
mean_number_of_floors = mean(Number.Of.Floors, na.rm = TRUE),
sd_number_of_floors = sd(Number.Of.Floors, na.rm = TRUE),
median_number_of_floors = median(Number.Of.Floors, na.rm = TRUE),
IQR_number_of_floors = IQR(Number.Of.Floors, na.rm = TRUE)
)
## number_observed mean_number_of_floors sd_number_of_floors
## 1 4774 1.661422 21.21494
## median_number_of_floors IQR_number_of_floors
## 1 1 1
The distribution of Number of Floors is right-skewed, so the appropriate statistics are the median and IQR. The median is 1 floor and the IQR is 1 floor.
Housing Units:
Number of household units for a given building
Discrete numeric variable
# Create a histogram for Housing Units
ggplot(data = permits) +
geom_histogram(
aes(x = Housing.Units),
binwidth = 1, center = 0.5,
color = "black", fill = "lightblue"
) +
# Scale the x axis from 0 to 10
scale_x_continuous(limits = c(0,10), breaks = seq(0, 10, 1)) +
# Add labels
labs(
title = "Distribution of Housing Units",
x = "Housing Units",
y = "count"
)
## Warning: Removed 14 rows containing non-finite outside the scale range
## (`stat_bin()`).
# Summary for Housing Units
permits |>
mutate(Housing.Units = as.numeric(Housing.Units)) |>
summarize(
number_observed = n(),
mean_housing_unit = mean(Housing.Units, na.rm = TRUE),
sd_housing_unit = sd(Housing.Units, na.rm = TRUE),
median_housing_unit = median(Housing.Units, na.rm = TRUE),
IQR_housing_unit = IQR(Housing.Units, na.rm = TRUE)
)
## number_observed mean_housing_unit sd_housing_unit median_housing_unit
## 1 4774 1.068582 3.031703 1
## IQR_housing_unit
## 1 0
The number of housing units is skewed so median and IQR are better measures than mean and SD. The median is 1 housing unit with an IQR of 0 housing units.
Issued Date:
Date on which the permit was issued
Discrete numeric variable (months are categorical ordinal)
# Create a histogram for Issued Date
permits |>
mutate(month_label = month(Issued.Date, label = TRUE)) |>
ggplot() +
geom_bar(aes(x = month_label)) +
labs(
x = "Month",
y = "Count",
title = "Permits issued per month"
)
# Summary for Issued Date
permits |>
mutate(month_label = month(Issued.Date, label = TRUE)) |>
group_by(month_label) |>
summarize(
num_observed = n()
)
## # A tibble: 10 × 2
## month_label num_observed
## <ord> <int>
## 1 Jan 827
## 2 Feb 801
## 3 Mar 934
## 4 Apr 828
## 5 May 526
## 6 Jun 456
## 7 Jul 221
## 8 Aug 115
## 9 Sep 61
## 10 Oct 5
The month of the issued date is a categorical ordinal variable, so we summarize it by count.
Some of the variables are dates, so they need to be handled differently than numbers. Creating a variable that stores the time between issuance and completion would help us answer the research question. For missing values, we can try out different policies of handling missing values (like dropping them or not) and see which one performs the best. We also noticed that there were no permits with Issuance Method of Online in the year 2026.
Wrangling: Based on what you discovered about your variables, prepare your dataset for analysis by performing any data wrangling that is appropriate for your project. This might include:
Then briefly summarize:
new_permits <-
permits |>
mutate(Housing.Units = as.numeric(Housing.Units)) |>
mutate(Number.Of.Floors = parse_number(Number.Of.Floors)) |>
mutate(Total.New.Add.SQFT = parse_number(Total.New.Add.SQFT)) |>
mutate(Time.Taken = as.numeric(difftime(Completed.Date, Issued.Date, units = "days")))
head(new_permits)
## Condominium Issuance.Method Status.Current Completed.Date Number.Of.Floors
## 1 No Permit Center Final 2026/08/19 1
## 2 No Permit Center Final 2026/09/28 1
## 3 No Permit Center Final 2026/06/02 1
## 4 No Permit Center Final 2026/09/04 1
## 5 No Permit Center Final 2026/09/25 1
## 6 No Permit Center Final 2026/10/02 1
## Housing.Units Issued.Date Total.New.Add.SQFT Time.Taken
## 1 1 2026/04/08 2062 133.00000
## 2 1 2026/03/06 5000 205.95833
## 3 1 2026/02/25 2446 96.95833
## 4 1 2026/01/23 1800 223.95833
## 5 1 2026/09/02 114 23.00000
## 6 2 2026/09/26 1500 6.00000
We filtered for where the total new add square feet is not null, the current status is “Final”, and the Calendar Year Issued is 2026. We also added a variable called Time.Taken which stores the number of days between the issued date and completed date
We will use the variables discussed above.
After filtering, there were no permits where the Issuance Method was “Online”. This may be an issue later on which we will need to consider. There may also be more data wrangling we need to do that we haven’t done yet.