Introduction

This code-through explores how R can be used to organize, combine, and visualize real-world data to identify trends in Arizona’s Extended Foster Care (EFC) Success Coaching program. EFC supports young adults, aged 17-21 years old, as they transition from foster care into adulthood, focusing on areas such as education, employment, housing stability, independent living skills, and supportive connections.

Using publicly available quarterly reports from the Arizona Department of Child Safety (DCS), this tutorial examines program participation and performance measures reported between March 2024 and June 2026. These reports provide information on the number of young adults served, participation in Administrative Reviews, transition planning, and other measures related to preparing young adults for adulthood.

For this code-through, I wanted to use data related to my work while practicing some of the skills we’ve learned in R. I’ll show how to organize quarterly statistics into datasets, combine them using left_join(), and create visualizations with ggplot2. I was interested in seeing how putting information from different reports together could make it easier to recognize trends in EFC performance.


Content Overview

This tutorial starts by creating separate datasets using enrollment and performance statistics from Arizona DCS quarterly reports. I’ll use left_join() from the dplyr package to combine the datasets using their shared reporting periods and show how to check for missing or unmatched data.

From there, I’ll use pivot_longer() to reorganize the data and ggplot2 to create graphs comparing performance measures over time. I’ll walk through the code and explain what the results show along the way.


Why You Should Care

When working with program data, information is often times spread across different reports, making it harder to see the bigger picture. This is especially relevant for programs like Extended Foster Care, where outcomes related to education, employment, housing, and transition planning are often connected. By combining and visualizing data in R, it can make it easier to identify trends, compare performance measures, and recognize areas of improvement. These skills can also be applied to other programs that use data to track services and outcomes.


Learning Objectives

By the end of this code-through, readers will be able to:

  1. Create and organize datasets in R using publicly reported quarterly performance statistics.

  2. Combine datasets using left_join() from the dplyr package and a shared reporting period.

  3. Identify and interpret missing or unmatched records when joining datasets.

  4. Restructure data using pivot_longer() from the tidyr package.

  5. Create and customize visualizations using ggplot2 to compare performance measures across reporting periods.

  6. Interpret trends in administrative data while recognizing limitations related to reporting periods and differences in the populations being measured.



Joining Arizona EFC Quarterly Data in R

For this example, I’ll use available information from Arizona DCS EFC quarterly reports between March 2024 and June 2026. I’ll create separate datasets for the number of young adults served and selected performance measures, including Administrative Review attendance, transition planning, and supportive connections. Although the information comes from the same reports, keeping it in separate datasets will help demonstrate how to join data using a shared variable, which is the reporting period.


Further Exposition

A dataset join combines information from two or more tables using a shared variable, also called a key. In R, the dplyr package includes functions for joining datasets, such as left_join(), inner_join(), and full_join(). The main difference between these functions is which rows they keep when combining the data.

For this tutorial, I’ll focus on left_join(), which keeps all the rows from the first dataset and adds matching information from the second dataset. If there isn’t a match, R fills in the missing information with N/A.

In the EFC datasets, the reporting period will be the shared variable, which allows us to match enrollment counts and performance measures from the same quarter. Before joining the datasets, it’s important to make sure the reporting periods are written the same way in both datasets. It is also important to make sure each quarter only appears once in each dataset, since duplicates could create extra rows when we combine them.


Basic Example

For this example, I used data from eight Arizona DCS EFC reports and manually entered the figures into R. I created two datasets: one with point in time enrollment counts for youth ages 17–20 served in the EFC Comprehensive Service Model and another with the percentage of young adults ages 18 and older with documented person-centered transition plans during Administrative Reviews. Even though these figures come from the original reports, I organized them into separate datasets to practice combining data using a shared reporting period.

# Create a dataset of youth served during each reporting period

efc_enrollment <- data.frame(
  quarter = c("Mar 2024", "Jun 2024", "Sep 2024",
              "Mar 2025", "Jun 2025", "Dec 2025",
              "Mar 2026", "Jun 2026"),
  youth_served = c(497, 709, 864, 1278, 1347, 1375, 1375, 1413)
)

# Display the dataset
efc_enrollment
##    quarter youth_served
## 1 Mar 2024          497
## 2 Jun 2024          709
## 3 Sep 2024          864
## 4 Mar 2025         1278
## 5 Jun 2025         1347
## 6 Dec 2025         1375
## 7 Mar 2026         1375
## 8 Jun 2026         1413

Next, a second dataset is created that shows the percentage of young adults ages 18 and older with documented person-centered transition plans during EFC Administrative Reviews. These percentages represent a different group than the enrollment counts, so it’s important to keep that distinction in mind when looking at the data.

# Create a dataset of documented transition plan percentages

efc_transition <- data.frame(
  quarter = c("Mar 2024", "Jun 2024", "Sep 2024",
              "Mar 2025", "Jun 2025", "Dec 2025",
              "Mar 2026", "Jun 2026"),
  transition_plan_pct = c(47, 79, 85, 80, 92, 93, 92, 94)
)

# Display the dataset
efc_transition
##    quarter transition_plan_pct
## 1 Mar 2024                  47
## 2 Jun 2024                  79
## 3 Sep 2024                  85
## 4 Mar 2025                  80
## 5 Jun 2025                  92
## 6 Dec 2025                  93
## 7 Mar 2026                  92
## 8 Jun 2026                  94

Now that both datasets have been created, I’ll use left_join() from the dplyr package to combine them. Since both datasets have a quarter column, R can use it to match the reporting periods and add the transition planning percentages to the enrollment data.Because each reporting period appears only once in both datasets, the combined dataset should still have eight rows. The difference is that we can now see the enrollment counts and transition planning percentages side by side.

# Combine enrollment and transition planning datasets

efc_combined <- dplyr::left_join(
  efc_enrollment,
  efc_transition,
  by = "quarter"
)

# Display the combined dataset
efc_combined
##    quarter youth_served transition_plan_pct
## 1 Mar 2024          497                  47
## 2 Jun 2024          709                  79
## 3 Sep 2024          864                  85
## 4 Mar 2025         1278                  80
## 5 Jun 2025         1347                  92
## 6 Dec 2025         1375                  93
## 7 Mar 2026         1375                  92
## 8 Jun 2026         1413                  94

After combining the datasets, I wanted to check whether any reporting periods were missing from the transition planning data. The anti_join() function from dplyr helps with this by showing which rows in the first dataset don’t have a match in the second dataset.I’ll use anti_join() to check whether every quarter in the enrollment dataset has a matching quarter in the transition planning dataset.

# Check for enrollment quarters without matching transition data

unmatched_quarters <- dplyr::anti_join(
  efc_enrollment,
  efc_transition,
  by = "quarter"
)

# Display unmatched reporting periods
unmatched_quarters
## [1] quarter      youth_served
## <0 rows> (or 0-length row.names)

The anti_join() returned zero rows means that every reporting period in the enrollment dataset has a match in the transition planning dataset. This is what is expected since both datasets use the same eight quarters. However, this doesn’t show whether the original reports contain missing or inaccurate information, but only confirms that the reporting periods match in this direction.


Advanced Examples

Now that the datasets are combined, I’ll add a couple of other EFC performance measures to look at more than just enrollment and transition planning. These include Administrative Review attendance and supportive connections. I’ll use left_join() to add these measures to the dataset and then use pivot_longer() and ggplot2 to organize and visualize the results. This will make it easier to see how the reported percentages have changed across the selected reporting periods.

# Create a dataset of additional EFC performance measures

efc_performance <- data.frame(
  quarter = c("Mar 2024", "Jun 2024", "Sep 2024",
              "Mar 2025", "Jun 2025", "Dec 2025",
              "Mar 2026", "Jun 2026"),
  review_attendance_pct = c(93, 86, 92, 93, 87, 83, 82, 82),
  supportive_connections_pct = c(88, 88, 90, 94, 95.5, 98, 98, 99)
)

# Display the dataset
efc_performance
##    quarter review_attendance_pct supportive_connections_pct
## 1 Mar 2024                    93                       88.0
## 2 Jun 2024                    86                       88.0
## 3 Sep 2024                    92                       90.0
## 4 Mar 2025                    93                       94.0
## 5 Jun 2025                    87                       95.5
## 6 Dec 2025                    83                       98.0
## 7 Mar 2026                    82                       98.0
## 8 Jun 2026                    82                       99.0


I’ll use left_join() again to add the Administrative Review attendance and supportive-connection percentages to the existing dataset. Since both datasets use quarter as their shared variable, R can match the reporting periods just like the Basic Example.This will give us one dataset with enrollment counts and all three performance measures, making it easier to work with the information together.

# Join additional performance measures to the combined dataset

efc_full <- dplyr::left_join(
  efc_combined,
  efc_performance,
  by = "quarter"
)

# Display the expanded dataset
efc_full
##    quarter youth_served transition_plan_pct review_attendance_pct
## 1 Mar 2024          497                  47                    93
## 2 Jun 2024          709                  79                    86
## 3 Sep 2024          864                  85                    92
## 4 Mar 2025         1278                  80                    93
## 5 Jun 2025         1347                  92                    87
## 6 Dec 2025         1375                  93                    83
## 7 Mar 2026         1375                  92                    82
## 8 Jun 2026         1413                  94                    82
##   supportive_connections_pct
## 1                       88.0
## 2                       88.0
## 3                       90.0
## 4                       94.0
## 5                       95.5
## 6                       98.0
## 7                       98.0
## 8                       99.0


The combined dataset is in wide format, with each performance measure in its own column, which makes it easy to read the information, but reorganizing it into long format will make it easier to create the graphs. Using pivot_longer() from the tidyr package will turn the three percentage columns into two new columns: measure, which identifies the performance measure, and percentage, which contains its value.

# Reshape EFC performance measures from wide to long format

efc_long <- tidyr::pivot_longer(
  efc_full,
  cols = c(transition_plan_pct,
           review_attendance_pct,
           supportive_connections_pct),
  names_to = "measure",
  values_to = "percentage"
)

# Display the reshaped dataset
efc_long
## # A tibble: 24 × 4
##    quarter  youth_served measure                    percentage
##    <chr>           <dbl> <chr>                           <dbl>
##  1 Mar 2024          497 transition_plan_pct                47
##  2 Mar 2024          497 review_attendance_pct              93
##  3 Mar 2024          497 supportive_connections_pct         88
##  4 Jun 2024          709 transition_plan_pct                79
##  5 Jun 2024          709 review_attendance_pct              86
##  6 Jun 2024          709 supportive_connections_pct         88
##  7 Sep 2024          864 transition_plan_pct                85
##  8 Sep 2024          864 review_attendance_pct              92
##  9 Sep 2024          864 supportive_connections_pct         90
## 10 Mar 2025         1278 transition_plan_pct                80
## # ℹ 14 more rows

After using pivot_longer(), the dataset now has twenty four rows instead of eight because each reporting period has three rows- one for each performance measure. We haven’t added any new information; we’ve just reorganized the existing data, which will make it easier to use ggplot2 and create separate graphs for each measure using facet_wrap().

With the data in long format, using ggplot2 will create a line graph showing how the EFC performance measures have changed over time. Also using facet_wrap() will give each measure its own graph instead of putting everything together in one, which will make it easier to see the trends for transition planning, Administrative Review attendance, and supportive connections.

This tutorial uses eight selected reporting periods between March 2024 and June 2026 rather than a complete quarterly time series. This means some quarters are missing, and changes between the reporting periods may have happened gradually. The performance measures may also represent different populations or use different reporting definitions. Because of this, the graphs should be used to identify trends in the reported data rather than determine whether the program caused those changes.

# Prepare the reporting periods in chronological order

efc_long$quarter <- factor(
  efc_long$quarter,
  levels = c("Mar 2024", "Jun 2024", "Sep 2024",
             "Mar 2025", "Jun 2025", "Dec 2025",
             "Mar 2026", "Jun 2026")
)

# Convert reporting periods into actual dates

efc_long$report_date <- as.Date(
  paste0("01 ", as.character(efc_long$quarter)),
  format = "%d %b %Y"
)

# Check that the dates were converted correctly
head(efc_long[, c("quarter", "report_date")])
## # A tibble: 6 × 2
##   quarter  report_date
##   <fct>    <date>     
## 1 Mar 2024 2024-03-01 
## 2 Mar 2024 2024-03-01 
## 3 Mar 2024 2024-03-01 
## 4 Jun 2024 2024-06-01 
## 5 Jun 2024 2024-06-01 
## 6 Jun 2024 2024-06-01
# Create a separate dataset for the visualization

efc_plot <- efc_long

# Convert performance measure names into readable labels

efc_plot$measure <- dplyr::recode(
  as.character(efc_plot$measure),
  "transition_plan_pct" = "Transition Planning",
  "review_attendance_pct" = "Administrative Review Attendance",
  "supportive_connections_pct" = "Supportive Connections"
)

# Create a line graph showing EFC performance trends

ggplot2::ggplot(
  efc_plot,
 ggplot2::aes(x = report_date, y = percentage, group = 1, color = measure)
)+
  ggplot2::geom_line(linewidth = 1) +
ggplot2::geom_point(size = 2.5) +
  ggplot2::facet_wrap(~measure, ncol = 1) +
  ggplot2::scale_color_manual(
  values = c(
    "Administrative Review Attendance" = "#247C83",
    "Supportive Connections" = "#75539A",
    "Transition Planning" = "#E58A45"
  ),
  guide = "none"
) +
  ggplot2::labs(
    title = "Arizona Extended Foster Care Performance Trends",
    subtitle = "Selected quarterly reports, March 2024–June 2026",
    x = "Reporting Period",
    y = "Reported Percentage (%)",
    caption = "Source: Arizona Department of Child Safety EFC Quarterly Reports"
  ) +
  ggplot2::theme_minimal() +
  ggplot2::theme(
    axis.text.x = ggplot2::element_text(
      angle = 45,
      hjust = 1
    ),
    plot.title = ggplot2::element_text(face = "bold")
  )

Looking at the graphs, transition planning increased from 47% in March 2024 to 94% in June 2026, while supportive connections increased from 88% to 99%. Administrative Review attendance followed a different pattern, decreasing from 93% to 82% during the same period. This makes it easier to see which reported measures increased and which decreased. Although these graphs help show changes over time, we need to be careful about how we interpret them. Some reporting periods are missing, and the performance measures may use different populations or definitions. We can identify trends in the reported data, but we can’t determine what caused those changes from these graphs alone.

Even though I used EFC data for this example, the same process could be useful for other programs that collect information across multiple reporting periods. Joining datasets and creating graphs can help make that information easier to organize, understand, and share.


Further Resources

Learn more about the R packages, data analysis techniques, and Arizona Extended Foster Care reports used in this tutorial through the following resources:



Works Cited

This code through references and cites the following sources:


  • Arizona Department of Child Safety. (2024). Extended Foster Care Comprehensive Service Model: Quarterly Report, March 2024. View report

  • Arizona Department of Child Safety. (2024). Extended Foster Care Comprehensive Service Model: Quarterly Report, June 2024. View report

  • Arizona Department of Child Safety. (2024). Extended Foster Care Comprehensive Service Model: Quarterly Report, September 2024. View report

  • Arizona Department of Child Safety. (2025). Extended Foster Care Comprehensive Service Model: Quarterly Report, March 2025. View report

  • Arizona Department of Child Safety. (2025). Extended Foster Care Comprehensive Service Model: Quarterly Report, June 2025. View report

  • Arizona Department of Child Safety. (2025). Extended Foster Care Comprehensive Service Model: Quarterly Report, December 2025. View report

  • Arizona Department of Child Safety. (2026). Extended Foster Care Comprehensive Service Model: Quarterly Report, March 2026. View report

  • Arizona Department of Child Safety. (2026). Extended Foster Care Comprehensive Service Model: Quarterly Report, June 2026. View report