The first steps to start the analysis was to find a data set. I chose the NREL historical data of alternative fuel stations to determine the build up of alternative electric vehicle fuel stations in the US. Once I determined the data set to use, I had to register for access to the data via their portal. After getting access to the data, I had to figure out how to use their API to download the data. Next was determining the fields that would be used in the analysis by studying the documentation and reading the data dictionary of the downloaded data. That work was spread over about three nights and next was researching how to get the intended graphs and data exploration. Fortunately, the data was not very sparse for the analysis I wanted to do. I took two nights getting familiar with the data and seeing what kinds of plots made sense and could be built. The remainder of my time was spent researching creating plots and improving graphs. I created eleven plots in total.
This is the Google registration function required to execute ggmap. Code will not function unless you have a Google API Key. This is my key.
First, pull map from OpenStreetMap
usa_map <- get_map(location = "USA", zoom = 4, source = "google")
## ā¹ <https://maps.googleapis.com/maps/api/staticmap?center=USA&zoom=4&size=640x640&scale=2&maptype=terrain&language=en-EN&key=xxx-rrxYKygXuR8DVgzRML1_jzg>
## ā¹ <https://maps.googleapis.com/maps/api/geocode/json?address=USA&key=xxx-rrxYKygXuR8DVgzRML1_jzg>
After pulling data from NREL API, I determined the column names and
created a list of columns that I potentially would use for my analysis.
I then repaired the column names by using the make.names
function to convert the white space to dots in column names. Then I
converted the data format in the Open.Date column.
electric_df <- alt_fuel_station_df
head(electric_df)
names(electric_df)
## [1] "Fuel Type Code"
## [2] "Station Name"
## [3] "Street Address"
## [4] "Intersection Directions"
## [5] "City"
## [6] "State"
## [7] "ZIP"
## [8] "Plus4"
## [9] "Station Phone"
## [10] "Status Code"
## [11] "Expected Date"
## [12] "Groups With Access Code"
## [13] "Access Days Time"
## [14] "Cards Accepted"
## [15] "BD Blends"
## [16] "NG Fill Type Code"
## [17] "NG PSI"
## [18] "EV Level1 EVSE Num"
## [19] "EV Level2 EVSE Num"
## [20] "EV DC Fast Count"
## [21] "EV Other Info"
## [22] "EV Network"
## [23] "EV Network Web"
## [24] "Geocode Status"
## [25] "Latitude"
## [26] "Longitude"
## [27] "Date Last Confirmed"
## [28] "ID"
## [29] "Updated At"
## [30] "Owner Type Code"
## [31] "Federal Agency ID"
## [32] "Federal Agency Name"
## [33] "Open Date"
## [34] "Hydrogen Status Link"
## [35] "NG Vehicle Class"
## [36] "LPG Primary"
## [37] "E85 Blender Pump"
## [38] "EV Connector Types"
## [39] "Country"
## [40] "Intersection Directions (French)"
## [41] "Access Days Time (French)"
## [42] "BD Blends (French)"
## [43] "Groups With Access Code (French)"
## [44] "Hydrogen Is Retail"
## [45] "Access Code"
## [46] "Access Detail Code"
## [47] "Federal Agency Code"
## [48] "Facility Type"
## [49] "CNG Dispenser Num"
## [50] "CNG On-Site Renewable Source"
## [51] "CNG Total Compression Capacity"
## [52] "CNG Storage Capacity"
## [53] "LNG On-Site Renewable Source"
## [54] "E85 Other Ethanol Blends"
## [55] "EV Pricing"
## [56] "EV Pricing (French)"
## [57] "LPG Nozzle Types"
## [58] "Hydrogen Pressures"
## [59] "Hydrogen Standards"
## [60] "CNG Fill Type Code"
## [61] "CNG PSI"
## [62] "CNG Vehicle Class"
## [63] "LNG Vehicle Class"
## [64] "EV On-Site Renewable Source"
## [65] "Restricted Access"
## [66] "RD Blends"
## [67] "RD Blends (French)"
## [68] "RD Blended with Biodiesel"
## [69] "RD Maximum Biodiesel Level"
## [70] "NPS Unit Name"
## [71] "CNG Station Sells Renewable Natural Gas"
## [72] "LNG Station Sells Renewable Natural Gas"
## [73] "Maximum Vehicle Class"
## [74] "EV Workplace Charging"
nrow(electric_df)
## [1] 84950
col_list <- c("Fuel Type Code", "ZIP", "State", "Status Code", "Expected Date","Open Date",
"EV Level1 EVSE Num", "EV Level2 EVSE Num", "EV DC Fast Count",
"EV Other Info", "EV Network", "EV Pricing", "Latitude", "Longitude",
"Facility Type", "Restricted Access", "Maximum Vehicle Class", "EV Workplace Charging")
electric_df |>
group_by(`Fuel Type Code`) |>
count()
sub_df <- electric_df |>
select(all_of(col_list)) |>
filter(`Fuel Type Code` == 'ELEC')
names(sub_df) <- make.names(names(sub_df), unique = TRUE)
sub_df$Open.Date <- ymd(sub_df$Open.Date)
sub_df$Expected.Date <- ymd(sub_df$Expected.Date)
head(sub_df)
Get summary of the data from the data frame to determine characteristics of the data like how many NAās exist in each applicable column, min and max value where applicable, and other data characteristics.
summary(sub_df)
## Fuel.Type.Code ZIP State Status.Code
## Length:73557 Length:73557 Length:73557 Length:73557
## Class :character Class :character Class :character Class :character
## Mode :character Mode :character Mode :character Mode :character
##
##
##
##
## Expected.Date Open.Date EV.Level1.EVSE.Num
## Min. :2019-08-16 Min. :1995-08-30 Min. : 1.00
## 1st Qu.:2023-06-15 1st Qu.:2020-02-27 1st Qu.: 1.00
## Median :2024-02-10 Median :2021-05-23 Median : 2.00
## Mean :2023-11-10 Mean :2020-11-25 Mean : 4.24
## 3rd Qu.:2024-05-16 3rd Qu.:2023-01-20 3rd Qu.: 3.00
## Max. :2026-12-01 Max. :2024-12-04 Max. :90.00
## NA's :68473 NA's :227 NA's :72835
## EV.Level2.EVSE.Num EV.DC.Fast.Count EV.Other.Info EV.Network
## Min. : 1.000 Min. : 1.00 Mode:logical Length:73557
## 1st Qu.: 2.000 1st Qu.: 1.00 NA's:73557 Class :character
## Median : 2.000 Median : 2.00 Mode :character
## Mean : 2.447 Mean : 4.06
## 3rd Qu.: 2.000 3rd Qu.: 6.00
## Max. :338.000 Max. :108.00
## NA's :10024 NA's :62540
## EV.Pricing Latitude Longitude Facility.Type
## Length:73557 Min. :18.01 Min. :-162.29 Length:73557
## Class :character 1st Qu.:34.04 1st Qu.:-117.92 Class :character
## Mode :character Median :38.55 Median : -92.10 Mode :character
## Mean :37.82 Mean : -96.48
## 3rd Qu.:41.52 3rd Qu.: -78.82
## Max. :64.85 Max. : -65.76
##
## Restricted.Access Maximum.Vehicle.Class EV.Workplace.Charging
## Mode :logical Length:73557 Mode :logical
## FALSE:10470 Class :character FALSE:72104
## TRUE :1156 Mode :character TRUE :1441
## NA's :61931 NA's :12
##
##
##
Outcome Strategy: I anticipate that EV Chargers will be more present in coastal states like California, New York, and Florida and fair number in Texas with a low number in the fly-over states. Also, expect mainly Level 1 and Level 2 chargers with a few Fast Chargers but expect to see growth in each type over time.
# Get count of fast chargers by state in descending order
sub_df |>
group_by(State) |>
summarize(total_fast_chargers = sum(EV.DC.Fast.Count, na.rm = TRUE)) |>
arrange(desc(total_fast_chargers))
# Summarize counts by state and charger type
state_charger_counts <- sub_df |>
group_by(State) |>
summarize(
Level1_Count = sum(EV.Level1.EVSE.Num, na.rm = TRUE),
Level2_Count = sum(EV.Level2.EVSE.Num, na.rm = TRUE),
Fast_Count = sum(EV.DC.Fast.Count, na.rm = TRUE)
) |>
pivot_longer(cols = c(Level1_Count, Level2_Count, Fast_Count),
names_to = "Charger_Type", values_to = "Count")
# Create the bar chart
ggplot(state_charger_counts, aes(x = State, y = Count, fill = Charger_Type)) +
geom_col(width = 0.7, position = position_dodge(width = 0.8)) +
labs(title = "Counts of Each Charger Type by State",
x = "State", y = "Charger Count",
fill = "Charger Type") +
theme_minimal() +
theme(axis.text.x = element_text(angle = 90, hjust = 1))
California dwarfs all other states in EV station infrastructure, with
New York and Texas in distant following. I would have imagined that
Texas would have been more on-par with California, but that was a
incorrect initial assumption.
Based on media presence, I would guess that Tesla to be the largest network.
# Group by EV.Network, count, and sort in descending order
sorted_data <- sub_df |>
group_by(EV.Network) |>
count() |>
arrange(desc(n))
# Identify the top 10 EV networks
top_10_networks <- sorted_data |>
arrange(desc(n)) |>
head(10) |>
pull(EV.Network)
# Filter data to include only the top 10 networks
sub_df_top10 <- sub_df |>
filter(EV.Network %in% top_10_networks) |>
group_by(EV.Network) |>
count() |>
arrange(desc(n))
# Create the bar plot
ggplot(sub_df_top10, aes(x = reorder(EV.Network, n), y = n, fill = EV.Network)) +
geom_col(width = 0.7, position = position_dodge(width = 0.8)) +
scale_fill_viridis_d(option = "D", guide = "none") +
labs(title = "Count of Top 10 EV Networks",
x = "EV Network", y = "Count") +
theme_minimal() +
theme(axis.text.x = element_text(angle = 45, hjust = 1))
Once again, Tesla is fifth in number of charging stations in US.
ChargePoint is the dominant player in this arena.
sub_df_no_CA <- sub_df |>
filter(State != "CA")
ggplot(sub_df, aes(x = State, fill = State)) +
geom_bar(width = 0.7) +
geom_hline(yintercept = 4000, linetype = "dashed", color = "red") + # Add horizontal line at 4000 count
scale_fill_viridis_d(option = "D", guide = "none") +
labs(title = "Distribution of Electric Vehicle Stations Across States",
x = "State", y = "Count") +
theme_minimal() +
theme(
axis.text.x = element_text(angle = 90, hjust = 1, vjust = 0.5),
plot.title = element_text(hjust = 0.5, size = 14, face = "bold"),
axis.title.x = element_text(margin = margin(t = 10)),
axis.title.y = element_text(margin = margin(r = 10))
) +
scale_x_discrete(expand = expansion(add = c(0.5, 0.5)))
# Excluding CA to see how other states compare
ggplot(sub_df_no_CA, aes(x = State, fill = State)) +
geom_bar(width = 0.7) +
geom_hline(yintercept = 4000, linetype = "dashed", color = "red") + # Add horizontal line at 4000 count
scale_fill_viridis_d(option = "D", guide = "none") +
labs(title = "Distribution of Electric Vehicle Stations Across States",
x = "State", y = "Count") +
theme_minimal() +
theme(
axis.text.x = element_text(angle = 90, hjust = 1, vjust = 0.5),
plot.title = element_text(hjust = 0.5, size = 14, face = "bold"),
axis.title.x = element_text(margin = margin(t = 10)),
axis.title.y = element_text(margin = margin(r = 10))
) +
scale_x_discrete(expand = expansion(add = c(0.5, 0.5)))
## Time series plot Observations: This plot shows the number of electric
vehicle stations in each state. California far exceeds all other states
in the number of charging stations, with NY in a distant second
place.
# Filter out rows with NA in Expected.Date and only dates after 2021
date_filtered_df <- sub_df |>
filter(!is.na(Expected.Date) & Expected.Date > as.Date("2021-12-31"))
# Group by Expected.Date and count occurrences
date_counts <- date_filtered_df |>
group_by(Expected.Date) |>
summarize(count = n())
# Time series plot
ggplot(date_counts, aes(x = Expected.Date, y = count)) +
geom_line(color = viridis::viridis(1)) +
geom_point(color = viridis::viridis(1)) +
labs(title = "Number of Expected EV Chargers Over Time",
x = "Expected Date", y = "Number of Chargers") +
theme_minimal() +
theme(
axis.text.x = element_text(angle = 45, hjust = 1),
plot.title = element_text(hjust = 0.5, size = 16, face = "bold"),
plot.subtitle = element_text(hjust = 0.5, size = 14)
) +
scale_x_date(date_breaks = "1 month", date_labels = "%b %Y", expand = expansion(mult = c(0.05, 0.05))) +
expand_limits(y = 0)
Showing the ratio of Level 1 to Level 2 Chargers in the US
ggplot(sub_df, aes(x = `EV.Level1.EVSE.Num`, y = `EV.Level2.EVSE.Num`)) +
geom_point(alpha = 0.6, color = "palegreen4", size = 2) +
geom_smooth(method = "lm", se = FALSE, color = "orangered2") +
labs(title = "Comparison of Level 1 and Level 2 EV Chargers",
subtitle = "Scatter plot with linear trend line",
x = "Number of Level 1 Chargers",
y = "Number of Level 2 Chargers") +
theme_minimal(base_size = 15) +
theme(
plot.title = element_text(hjust = 0.5, size = 16, face = "bold"),
plot.subtitle = element_text(hjust = 0.5, size = 14),
axis.title.x = element_text(margin = margin(t = 10)),
axis.title.y = element_text(margin = margin(r = 10))
)
## `geom_smooth()` using formula = 'y ~ x'
Observations: This scatter plot helps to understand the relationship between the number of Level 1 and Level 2 chargers at each station.
# Get the map data for the USA
state_map <- map_data("state")
# Plot the map with EV station locations
ev1_map <- ggplot() +
geom_polygon(data = state_map, aes(x = long, y = lat, group = group), fill = "white", color = "black") +
geom_point(data = sub_df, aes(x = Longitude,
y = Latitude, color = "Level 1",
size = `EV.Level1.EVSE.Num`),
alpha = 0.5) +
labs(title = "Geographical Distribution of EV Level 1 Chargers in the USA",
x = "Longitude", y = "Latitude",
color = "Charger Type", size = "Count") +
theme_minimal() +
theme(plot.margin = margin(1, 1, 1, 1))
print(ev1_map)
ggsave("ev1_map.png")
## Saving 7 x 5 in image
Observations: This scatter plot overlays the counts of Level 1 chargers at each station.
The ggmap plots may give error message when they are first run. This is some kind of API issue which re-executing the code seems to be the only thing I found to work.
ev_ggmap <- ggmap(usa_map, extent='panel', maprange = TRUE) +
geom_point(data = sub_df, aes(x = Longitude, y = Latitude,
color = `EV.Level2.EVSE.Num`, size = `EV.Level2.EVSE.Num`), alpha = 0.3, shape = 1) + # Plot Level 2 chargers second
geom_point(data = sub_df, aes(x = Longitude, y = Latitude,
color = `EV.DC.Fast.Count`, size = `EV.DC.Fast.Count`), alpha = 0.05, shape = 17) + # Plot DC Fast chargers first
geom_point(data = sub_df, aes(x = Longitude, y = Latitude,
color = `EV.Level1.EVSE.Num`, size = `EV.Level1.EVSE.Num`), alpha = 0.9, shape = 15) + # Plot Level 1 chargers last
scale_color_gradient(low = "dodgerblue2", high = "red1") +
labs(title = "Geographical Distribution of EV Chargers in the USA",
x = "Longitude", y = "Latitude",
color = "Charger Count", size = "Charger Count") +
theme_minimal() +
theme(plot.margin = margin(1, 1, 1, 1))
## Scale for x is already present.
## Adding another scale for x, which will replace the existing scale.
## Scale for y is already present.
## Adding another scale for y, which will replace the existing scale.
print(ev_ggmap)
ggsave("ev_ggmap.png")
## Saving 7 x 5 in image
sub_df_long <- sub_df |>
pivot_longer(cols = c("EV.Level1.EVSE.Num", "EV.Level2.EVSE.Num", "EV.DC.Fast.Count"),
names_to = "Charger.Type", values_to = "Charger.Count") |>
filter(Charger.Count > 0)
alpha_values <- c("EV.Level1.EVSE.Num" = 0.9,
"EV.Level2.EVSE.Num" = 0.02,
"EV.DC.Fast.Count" = 0.01)
ev_comp_ggmap <- ggmap(usa_map, extent='panel', maprange = TRUE) +
geom_point(data = sub_df_long, aes(x = Longitude, y = Latitude,
color = Charger.Type, size = Charger.Count), alpha = 0.5) +
scale_color_manual(values = c("EV.Level1.EVSE.Num" = "dodgerblue3",
"EV.Level2.EVSE.Num" = "darkolivegreen4",
"EV.DC.Fast.Count" = "firebrick3"
)
) +
scale_alpha_manual(values = alpha_values) +
scale_size_continuous(range = c(1, 5)) +
labs(title = "Geographical Distribution of EV Chargers in the USA",
x = "Longitude", y = "Latitude",
color = "Charger Type", size = "Charger Count") +
theme_minimal() +
theme(plot.margin = margin(1, 1, 1, 1))
## Scale for x is already present.
## Adding another scale for x, which will replace the existing scale.
## Scale for y is already present.
## Adding another scale for y, which will replace the existing scale.
print(ev_comp_ggmap)
png("ev_comp_ggmap.png")
Even though all charger types are present in this plot, the level 1 chargers bleed out with the majority of level 2 chargers crowding out the representation of level 1 charging stations.
ev_level1_map <- ggmap(usa_map, extent='normal', maprange = TRUE) +
geom_point(data = sub_df, aes(x = Longitude, y = Latitude, color = `EV.Level1.EVSE.Num`, size = `EV.Level1.EVSE.Num`), alpha = 0.6) +
scale_color_gradient(low = "dodgerblue2", high = "firebrick1") +
labs(title = "Geographical Distribution of EV Fast Charger in the USA",
x = "Longitude", y = "Latitude",
color = "Charger Count", size = "Charger Count") +
theme_minimal() +
theme(plot.margin = margin(1, 1, 1, 1))
print(ev_level1_map)
ggsave("ev_level1_map.png")
## Saving 7 x 5 in image
The sparsity and concentration of Level 1 chargers get drowned out in the other plots. They are available in certain major metropolitan areas, but do not have any noticeable presence beyond that.
sub_df_long <- sub_df |>
mutate(Open.Date = as.Date(Open.Date, format = "%Y-%m-%d")) |>
mutate(Year = year(Open.Date)) |>
filter(Open.Date >= as.Date("2010-01-01")) |>
pivot_longer(cols = c(`EV.Level1.EVSE.Num`, `EV.Level2.EVSE.Num`, `EV.DC.Fast.Count`),
names_to = "ChargerType", values_to = "Count") |>
mutate(ChargerType = recode(ChargerType,
`EV.Level1.EVSE.Num` = "Level 1",
`EV.Level2.EVSE.Num` = "Level 2",
`EV.DC.Fast.Count` = "DC Fast")) |>
select(c("Year", "ZIP", "ChargerType", "Count")) |>
group_by(Year, ChargerType) |> # Group by both Year and ChargerType
summarize(TotalCount = sum(Count, na.rm = TRUE)) # Summarize by summing the Count
## `summarise()` has grouped output by 'Year'. You can override using the
## `.groups` argument.
# Plot the cumulative count of each charger type over time
ggplot(sub_df_long, aes(x = Year, y = TotalCount, color = ChargerType)) +
geom_line() +
facet_wrap(~ ChargerType, scales = "free_y") + # Facet by ChargerType with independent y-axes
scale_color_manual(values = c("Level 1" = "dodgerblue2", "Level 2" = "palegreen3", "DC Fast" = "rosybrown4")) +
labs(title = "Cumulative Count of EV Chargers Over Time",
x = "Year", y = "Cumulative Count", color = "Charger Type") +
theme_minimal() +
theme(plot.title = element_text(hjust = 0.5))
Observations: This plot shows the cumulative number of each charger type over time. We can observe how the availability of Level 1, Level 2, and DC Fast chargers has changed over time.
ggplot(sub_df, aes(x = `EV.DC.Fast.Count`)) +
geom_histogram(binwidth = 1) +
labs(title = "Distribution of EV DC Fast Chargers",
x = "Number of DC Fast Chargers", y = "Count") +
theme_minimal()
Most locations with a DC Fast Charger only have 1 charger. The number drops significantly as the number of chargers available increases.