This code-through explores how R can be used to organize, combine, and visualize real-world data to identify trends in Arizona’s Extended Foster Care (EFC) Success Coaching program. EFC supports young adults, aged 17-21 years old, as they transition from foster care into adulthood, focusing on areas such as education, employment, housing stability, independent living skills, and supportive connections.
Using publicly available quarterly reports from the Arizona Department of Child Safety (DCS), this tutorial examines program participation and performance measures reported between March 2024 and June 2026. These reports provide information on the number of young adults served, participation in Administrative Reviews, transition planning, and other measures related to preparing young adults for adulthood.
For this code-through, I wanted to use data related to my work while practicing some of the skills we’ve learned in R. I’ll show how to organize quarterly statistics into datasets, combine them using left_join(), and create visualizations with ggplot2. I was interested in seeing how putting information from different reports together could make it easier to recognize trends in EFC performance.
This tutorial starts by creating separate datasets using enrollment and performance statistics from Arizona DCS quarterly reports. I’ll use left_join() from the dplyr package to combine the datasets using their shared reporting periods and show how to check for missing or unmatched data.
From there, I’ll use pivot_longer() to reorganize the data and ggplot2 to create graphs comparing performance measures over time. I’ll walk through the code and explain what the results show along the way.
When working with program data, information is often times spread across different reports, making it harder to see the bigger picture. This is especially relevant for programs like Extended Foster Care, where outcomes related to education, employment, housing, and transition planning are often connected. By combining and visualizing data in R, it can make it easier to identify trends, compare performance measures, and recognize areas of improvement. These skills can also be applied to other programs that use data to track services and outcomes.
By the end of this code-through, readers will be able to:
Create and organize datasets in R using publicly reported quarterly performance statistics.
Combine datasets using left_join() from the dplyr package and a shared reporting period.
Identify and interpret missing or unmatched records when joining datasets.
Restructure data using pivot_longer() from the tidyr package.
Create and customize visualizations using ggplot2 to compare performance measures across reporting periods.
Interpret trends in administrative data while recognizing limitations related to reporting periods and differences in the populations being measured.
For this example, I’ll use available information from Arizona DCS EFC quarterly reports between March 2024 and June 2026. I’ll create separate datasets for the number of young adults served and selected performance measures, including Administrative Review attendance, transition planning, and supportive connections. Although the information comes from the same reports, keeping it in separate datasets will help demonstrate how to join data using a shared variable, which is the reporting period.
A dataset join combines information from two or more tables using a shared variable, also called a key. In R, the dplyr package includes functions for joining datasets, such as left_join(), inner_join(), and full_join(). The main difference between these functions is which rows they keep when combining the data.
For this tutorial, I’ll focus on left_join(), which keeps all the rows from the first dataset and adds matching information from the second dataset. If there isn’t a match, R fills in the missing information with N/A.
In the EFC datasets, the reporting period will be the shared variable, which allows us to match enrollment counts and performance measures from the same quarter. Before joining the datasets, it’s important to make sure the reporting periods are written the same way in both datasets. It is also important to make sure each quarter only appears once in each dataset, since duplicates could create extra rows when we combine them.
For this example, I used data from eight Arizona DCS EFC reports and manually entered the figures into R. I created two datasets: one with point in time enrollment counts for youth ages 17–20 served in the EFC Comprehensive Service Model and another with the percentage of young adults ages 18 and older with documented person-centered transition plans during Administrative Reviews. Even though these figures come from the original reports, I organized them into separate datasets to practice combining data using a shared reporting period.
# Create a dataset of youth served during each reporting period
efc_enrollment <- data.frame(
quarter = c("Mar 2024", "Jun 2024", "Sep 2024",
"Mar 2025", "Jun 2025", "Dec 2025",
"Mar 2026", "Jun 2026"),
youth_served = c(497, 709, 864, 1278, 1347, 1375, 1375, 1413)
)
# Display the dataset
efc_enrollment## quarter youth_served
## 1 Mar 2024 497
## 2 Jun 2024 709
## 3 Sep 2024 864
## 4 Mar 2025 1278
## 5 Jun 2025 1347
## 6 Dec 2025 1375
## 7 Mar 2026 1375
## 8 Jun 2026 1413
Next, a second dataset is created that shows the percentage of young adults ages 18 and older with documented person-centered transition plans during EFC Administrative Reviews. These percentages represent a different group than the enrollment counts, so it’s important to keep that distinction in mind when looking at the data.
# Create a dataset of documented transition plan percentages
efc_transition <- data.frame(
quarter = c("Mar 2024", "Jun 2024", "Sep 2024",
"Mar 2025", "Jun 2025", "Dec 2025",
"Mar 2026", "Jun 2026"),
transition_plan_pct = c(47, 79, 85, 80, 92, 93, 92, 94)
)
# Display the dataset
efc_transition## quarter transition_plan_pct
## 1 Mar 2024 47
## 2 Jun 2024 79
## 3 Sep 2024 85
## 4 Mar 2025 80
## 5 Jun 2025 92
## 6 Dec 2025 93
## 7 Mar 2026 92
## 8 Jun 2026 94
Now that both datasets have been created, I’ll use left_join() from the dplyr package to combine them. Since both datasets have a quarter column, R can use it to match the reporting periods and add the transition planning percentages to the enrollment data.Because each reporting period appears only once in both datasets, the combined dataset should still have eight rows. The difference is that we can now see the enrollment counts and transition planning percentages side by side.
# Combine enrollment and transition planning datasets
efc_combined <- dplyr::left_join(
efc_enrollment,
efc_transition,
by = "quarter"
)
# Display the combined dataset
efc_combined## quarter youth_served transition_plan_pct
## 1 Mar 2024 497 47
## 2 Jun 2024 709 79
## 3 Sep 2024 864 85
## 4 Mar 2025 1278 80
## 5 Jun 2025 1347 92
## 6 Dec 2025 1375 93
## 7 Mar 2026 1375 92
## 8 Jun 2026 1413 94
After combining the datasets, I wanted to check whether any reporting periods were missing from the transition planning data. The anti_join() function from dplyr helps with this by showing which rows in the first dataset don’t have a match in the second dataset.I’ll use anti_join() to check whether every quarter in the enrollment dataset has a matching quarter in the transition planning dataset.
# Check for enrollment quarters without matching transition data
unmatched_quarters <- dplyr::anti_join(
efc_enrollment,
efc_transition,
by = "quarter"
)
# Display unmatched reporting periods
unmatched_quarters## [1] quarter youth_served
## <0 rows> (or 0-length row.names)
The anti_join() returned zero rows means that every reporting period in the enrollment dataset has a match in the transition planning dataset. This is what is expected since both datasets use the same eight quarters. However, this doesn’t show whether the original reports contain missing or inaccurate information, but only confirms that the reporting periods match in this direction.
Now that the datasets are combined, I’ll add a couple of other EFC performance measures to look at more than just enrollment and transition planning. These include Administrative Review attendance and supportive connections. I’ll use left_join() to add these measures to the dataset and then use pivot_longer() and ggplot2 to organize and visualize the results. This will make it easier to see how the reported percentages have changed across the selected reporting periods.
# Create a dataset of additional EFC performance measures
efc_performance <- data.frame(
quarter = c("Mar 2024", "Jun 2024", "Sep 2024",
"Mar 2025", "Jun 2025", "Dec 2025",
"Mar 2026", "Jun 2026"),
review_attendance_pct = c(93, 86, 92, 93, 87, 83, 82, 82),
supportive_connections_pct = c(88, 88, 90, 94, 95.5, 98, 98, 99)
)
# Display the dataset
efc_performance## quarter review_attendance_pct supportive_connections_pct
## 1 Mar 2024 93 88.0
## 2 Jun 2024 86 88.0
## 3 Sep 2024 92 90.0
## 4 Mar 2025 93 94.0
## 5 Jun 2025 87 95.5
## 6 Dec 2025 83 98.0
## 7 Mar 2026 82 98.0
## 8 Jun 2026 82 99.0
I’ll use left_join() again to add the Administrative Review attendance and supportive-connection percentages to the existing dataset. Since both datasets use quarter as their shared variable, R can match the reporting periods just like the Basic Example.This will give us one dataset with enrollment counts and all three performance measures, making it easier to work with the information together.
# Join additional performance measures to the combined dataset
efc_full <- dplyr::left_join(
efc_combined,
efc_performance,
by = "quarter"
)
# Display the expanded dataset
efc_full## quarter youth_served transition_plan_pct review_attendance_pct
## 1 Mar 2024 497 47 93
## 2 Jun 2024 709 79 86
## 3 Sep 2024 864 85 92
## 4 Mar 2025 1278 80 93
## 5 Jun 2025 1347 92 87
## 6 Dec 2025 1375 93 83
## 7 Mar 2026 1375 92 82
## 8 Jun 2026 1413 94 82
## supportive_connections_pct
## 1 88.0
## 2 88.0
## 3 90.0
## 4 94.0
## 5 95.5
## 6 98.0
## 7 98.0
## 8 99.0
The combined dataset is in wide format, with each performance measure in its own column, which makes it easy to read the information, but reorganizing it into long format will make it easier to create the graphs. Using pivot_longer() from the tidyr package will turn the three percentage columns into two new columns: measure, which identifies the performance measure, and percentage, which contains its value.
# Reshape EFC performance measures from wide to long format
efc_long <- tidyr::pivot_longer(
efc_full,
cols = c(transition_plan_pct,
review_attendance_pct,
supportive_connections_pct),
names_to = "measure",
values_to = "percentage"
)
# Display the reshaped dataset
efc_long## # A tibble: 24 × 4
## quarter youth_served measure percentage
## <chr> <dbl> <chr> <dbl>
## 1 Mar 2024 497 transition_plan_pct 47
## 2 Mar 2024 497 review_attendance_pct 93
## 3 Mar 2024 497 supportive_connections_pct 88
## 4 Jun 2024 709 transition_plan_pct 79
## 5 Jun 2024 709 review_attendance_pct 86
## 6 Jun 2024 709 supportive_connections_pct 88
## 7 Sep 2024 864 transition_plan_pct 85
## 8 Sep 2024 864 review_attendance_pct 92
## 9 Sep 2024 864 supportive_connections_pct 90
## 10 Mar 2025 1278 transition_plan_pct 80
## # ℹ 14 more rows
After using pivot_longer(), the dataset now has twenty four rows instead of eight because each reporting period has three rows- one for each performance measure. We haven’t added any new information; we’ve just reorganized the existing data, which will make it easier to use ggplot2 and create separate graphs for each measure using facet_wrap().
With the data in long format, using ggplot2 will create a line graph showing how the EFC performance measures have changed over time. Also using facet_wrap() will give each measure its own graph instead of putting everything together in one, which will make it easier to see the trends for transition planning, Administrative Review attendance, and supportive connections.
This tutorial uses eight selected reporting periods between March 2024 and June 2026 rather than a complete quarterly time series. This means some quarters are missing, and changes between the reporting periods may have happened gradually. The performance measures may also represent different populations or use different reporting definitions. Because of this, the graphs should be used to identify trends in the reported data rather than determine whether the program caused those changes.
# Prepare the reporting periods in chronological order
efc_long$quarter <- factor(
efc_long$quarter,
levels = c("Mar 2024", "Jun 2024", "Sep 2024",
"Mar 2025", "Jun 2025", "Dec 2025",
"Mar 2026", "Jun 2026")
)
# Convert reporting periods into actual dates
efc_long$report_date <- as.Date(
paste0("01 ", as.character(efc_long$quarter)),
format = "%d %b %Y"
)
# Check that the dates were converted correctly
head(efc_long[, c("quarter", "report_date")])## # A tibble: 6 × 2
## quarter report_date
## <fct> <date>
## 1 Mar 2024 2024-03-01
## 2 Mar 2024 2024-03-01
## 3 Mar 2024 2024-03-01
## 4 Jun 2024 2024-06-01
## 5 Jun 2024 2024-06-01
## 6 Jun 2024 2024-06-01
# Create a separate dataset for the visualization
efc_plot <- efc_long
# Convert performance measure names into readable labels
efc_plot$measure <- dplyr::recode(
as.character(efc_plot$measure),
"transition_plan_pct" = "Transition Planning",
"review_attendance_pct" = "Administrative Review Attendance",
"supportive_connections_pct" = "Supportive Connections"
)
# Create a line graph showing EFC performance trends
ggplot2::ggplot(
efc_plot,
ggplot2::aes(x = report_date, y = percentage, group = 1, color = measure)
)+
ggplot2::geom_line(linewidth = 1) +
ggplot2::geom_point(size = 2.5) +
ggplot2::facet_wrap(~measure, ncol = 1) +
ggplot2::scale_color_manual(
values = c(
"Administrative Review Attendance" = "#247C83",
"Supportive Connections" = "#75539A",
"Transition Planning" = "#E58A45"
),
guide = "none"
) +
ggplot2::labs(
title = "Arizona Extended Foster Care Performance Trends",
subtitle = "Selected quarterly reports, March 2024–June 2026",
x = "Reporting Period",
y = "Reported Percentage (%)",
caption = "Source: Arizona Department of Child Safety EFC Quarterly Reports"
) +
ggplot2::theme_minimal() +
ggplot2::theme(
axis.text.x = ggplot2::element_text(
angle = 45,
hjust = 1
),
plot.title = ggplot2::element_text(face = "bold")
)Looking at the graphs, transition planning increased from 47% in March 2024 to 94% in June 2026, while supportive connections increased from 88% to 99%. Administrative Review attendance followed a different pattern, decreasing from 93% to 82% during the same period. This makes it easier to see which reported measures increased and which decreased. Although these graphs help show changes over time, we need to be careful about how we interpret them. Some reporting periods are missing, and the performance measures may use different populations or definitions. We can identify trends in the reported data, but we can’t determine what caused those changes from these graphs alone.
Even though I used EFC data for this example, the same process could be useful for other programs that collect information across multiple reporting periods. Joining datasets and creating graphs can help make that information easier to organize, understand, and share.
Learn more about the R packages, data analysis techniques, and Arizona Extended Foster Care reports used in this tutorial through the following resources:
Resource I: Mutating Joins — dplyr— Explains how left_join() combines datasets using shared variables.
Resource II: Filtering Joins — dplyr— Demonstrates how anti_join() identifies records without matching observations.
Resource III: Pivot Data from Wide to Long — tidyr— Explains how to restructure data from wide to long format for analysis and visualization.
Resource IV: ggplot2 Documentation— Provides examples of creating and customizing graphs in R, including line graphs and faceting
Resource V: Arizona Department of Child Safety Reports— Provides access to the original EFC quarterly reports used to construct the datasets in this tutorial.
This code through references and cites the following sources:
Arizona Department of Child Safety. (2024). Extended Foster Care Comprehensive Service Model: Quarterly Report, March 2024. View report
Arizona Department of Child Safety. (2024). Extended Foster Care Comprehensive Service Model: Quarterly Report, June 2024. View report
Arizona Department of Child Safety. (2024). Extended Foster Care Comprehensive Service Model: Quarterly Report, September 2024. View report
Arizona Department of Child Safety. (2025). Extended Foster Care Comprehensive Service Model: Quarterly Report, March 2025. View report
Arizona Department of Child Safety. (2025). Extended Foster Care Comprehensive Service Model: Quarterly Report, June 2025. View report
Arizona Department of Child Safety. (2025). Extended Foster Care Comprehensive Service Model: Quarterly Report, December 2025. View report
Arizona Department of Child Safety. (2026). Extended Foster Care Comprehensive Service Model: Quarterly Report, March 2026. View report
Arizona Department of Child Safety. (2026). Extended Foster Care Comprehensive Service Model: Quarterly Report, June 2026. View report