Introduction

Learning a new language is not as simple as studying words, it requires progress in multiple areas. Simply tracking the amount of time that someone has studied does not accurately show their progress or aptitude. It is important to track and identify vocabulary mastered, words read, listening hours, speaking hours, and days studied. Collecting and organizing this data can help identify current standing, measure progress, and potentially forecast future proficiency.

Content Overview

This code through will demonstrate how R can be used to analyze language learning data. As an example, we will use the Hungarian language and track vocabulary mastery, words read, listening hours, speaking hours and days studied. These measures will be used to identify and visualize current status, progress over time, potential future progress and areas that need additional attention.

Why You Should Care

Language learning is not a simple task, especially self directed language learning. It is important to follow a plan and track progress over time, however, it can be difficult to identify areas for improvement and current progress without proper data. For example, you might feel like you are reading and understanding a lot, but when you go to have a conversation, you cannot think of the words to use or how to use them. This can be defeating and halt progress. Tracking language learning data can help organize progress, identify areas that need improvement, and motivate a student by showing how much they have progressed over time and what their potential future progress could look like.

Learning Objectives

Specifically, you’ll learn how to:

  1. Organize different measures of language learning progress into a dataset in R.
  2. Analyze changes in vocabulary, reading, listening, speaking, and study consistency over time.
  3. Identify areas of strength and areas that may need additional improvement.
  4. Visualize progress and use trends in the data to estimate potential future progress.

Tracking Language Learning Progress with R

In order to track language learning progress, we first need to create a dataset that measures progress consistently over time. For this example, we will track five areas of Hungarian language learning: vocabulary mastered, words read, listening hours, speaking hours, and days studied.

Measuring Language Learning Progress

Language learning progress can be measured in multiple ways. For this example, six months of learning data will be used to demonstrate how progress can be tracked over time in a simple way. The dataset includes cumulative vocabulary mastered, words read, listening hours, and speaking hours, along with the number of days studied each month.

I compare my language learning activity data to estimated CEFR proficiency benchmarks to see whether all variables are consistent with the same general proficiency level. The CEFR ranges from A1, representing a beginner language user, through C2, representing a highly proficient language user. This comparison does not denote a CEFR level exactly, but it can show whether each area of language learning falls within the general range expected for that level.

Creating the Language Learning Dataset

First, we will make a dataset containing six months of language learning activity. Each row represents one month of study, while the columns represent the different measures that are being tracked.

language.progress <- data.frame(
  month = c(1, 2, 3, 4, 5, 6),
  vocabulary = c(300, 550, 800, 1050, 1300, 1600),
  words.read = c(5000, 12000, 20000, 29000, 39000, 50835),
  listening.hours = c(3, 6, 9, 12, 16, 20),
  speaking.hours = c(0.5, 1, 1.5, 2, 2.5, 3),
  days.studied = c(12, 14, 15, 16, 18, 15)
)

language.progress


Analyzing Progress Over Time

Vocabulary Progress

Now that we have simple data, we can see how language learning has changed over the six month time frame. The following graph shows growth in vocabulary over time.

plot(language.progress$month,
     language.progress$vocabulary,
     type = "o",
     xlab = "Month",
     ylab = "Vocabulary Mastered",
     main = "Vocabulary Progress Over Six Months",
     ylim = c(0, 1600),
     yaxt = "n")

axis(2, at = c(0, 400, 800, 1200, 1600))


Now that we can visualize vocabulary over time, we can compare vocabulary with estimated CEFR benchmarks. These ranges do not represent an official CEFR assesment, but provide context for the data.

vocabulary.current <- 1600

if (vocabulary.current < 500) {
  vocabulary.level <- "Below A1"
} else if (vocabulary.current < 1000) {
  vocabulary.level <- "A1"
} else if (vocabulary.current < 2000) {
  vocabulary.level <- "A2"
} else if (vocabulary.current < 5000) {
  vocabulary.level <- "B1"
} else if (vocabulary.current < 8000) {
  vocabulary.level <- "B2"
} else if (vocabulary.current < 15000) {
  vocabulary.level <- "C1"
} else {
  vocabulary.level <- "C2"
}

vocabulary.level
## [1] "A2"


Reading Progress

Another way to measure language learning is through number of words read. This identifies the cumulative number of Hungarian words read during the six month period.

Total number of words read can be compared with the estimated CEFR reading benchmarks. This will help us to understand if reading exposure is progressing at the same rate as vocabulary.

words.read.current <- 50835
reading.level <- "A1"

reading.level
## [1] "A1"


Listening Progress

Listening exposure measures time spent hearing the language, which can be tracked to determine if listening practice is progressing at the same rate as other variables. Listening hours can be compared with estimated CEFR benchmarks.

listening.current <- 20
listening.level <- "Below A1"

listening.level
## [1] "Below A1"


Speaking Hours

Speaking hours represents active language use. Speaking hours can develop differently and much later than other language learning variables. This can identify if active language use is keepinig pace with the other areas of learning. Speaking hours can be compared with estimated CEFR benchmarks to identify if active language use is progressing at a similar rate to the other measures.

speaking.current <- 3
speaking.level <- "Below A1"

speaking.level
## [1] "Below A1"


Study Consistency

Days of language learning study can help to provide context for the other measurements. This variable will not be measured against a CEFR benchmark. This will help identify consistency of practice during the month.

plot(language.progress$month,
     language.progress$days.studied,
     type = "o",
     xlab = "Month",
     ylab = "Days Studied",
     main = "Study Consistency Over Six Months",
     ylim = c(0, 20),
     yaxt = "n")

axis(2, at = c(0, 5, 10, 15, 20))


Comparing Progress Across Measures

Vocabulary, words read, listening hours, and speaking hours are measured differently and cannot be direclty compared on the same scale. Instead, we can convert each measure to a percentage of the estimated benchmark to compare how each area has progressed over time.

progress.percent <- data.frame(
  Month = language.progress$month,
  Vocabulary = (language.progress$vocabulary / 2000) * 100,
  Words.Read = (language.progress$words.read / 100000) * 100,
  Listening = (language.progress$listening.hours / 150) * 100,
  Speaking = (language.progress$speaking.hours / 30) * 100
)

progress.percent
matplot(
  progress.percent$Month,
  progress.percent[, 2:5],
  type = "o",
  lty = 1,
  pch = 1:4,
  col = c("#4472C4", "#70AD47", "#ED7D31", "#8064A2"),
  xlab = "Month",
  ylab = "Percent of A2 Benchmark",
  main = "Language Learning Progress Over Six Months",
  ylim = c(0, 100),
  yaxt = "n"
)

axis(
  2,
  at = c(0, 20, 40, 60, 80, 100),
  labels = c("0%", "20%", "40%", "60%", "80%", "100%")
)

legend(
  "topleft",
  legend = c("Vocabulary", "Words Read", "Listening", "Speaking"),
  lty = 1,
  pch = 1:4,
    col = c("#4472C4", "#70AD47", "#ED7D31", "#8064A2")
)


Overall Progress Estimate

Visualizing multiple measures together can provide a clear picture toward an estimated overall proficiency. Vocabulary falls withing the A2 range, words read are closer to A1 and listening and speaking remain below A1. We may be able to describe the overall estimate to be around A1 without taking a certified CEFR assessment. This can help tailor the language learning plan prior to an assessment.

overall.progress <- data.frame(
  Measure = c("Vocabulary", "Words Read", "Listening",
              "Speaking", "Overall Estimate"),
  Estimated.Level = c(vocabulary.level, reading.level,
                      listening.level, speaking.level,
                      "Around A1")
)

overall.progress


Planning Future Progress

The above information can help to tailor the language learning plan. R can be used to estimate how long it may take to reach learning goals. Firt, we can compare current measurements with target values and monthly progress to calculate the approximate number of months needed for each variable to achieve A2 in six months. Then, we can identify how to alter the plan to reach A2 level across all variables in three months. These estimates do not guarantee achievement of official CEFR levels.

Projecting Progress Using Current Study Plan

a2.plan <- data.frame(
  Measure = c("Vocabulary", "Words Read", "Listening", "Speaking"),
  Current = c(1600, 50835, 20, 3),
  A2.Target = c(2000, 100000, 150, 30),
  Monthly.Goal = c(300, 20000, 20, 5)
)

a2.plan$Months.To.A2 <- ceiling(
  (a2.plan$A2.Target - a2.plan$Current) / a2.plan$Monthly.Goal
)

a2.plan
plot(
  a2.plan$Months.To.A2,
  1:4,
  pch = 19,
  cex = 1.5,
  xlim = c(0, 8),
  ylim = c(0.5, 4.5),
  xlab = "Months From Now",
  ylab = "",
  main = "Projected Timeline Toward A2 Learning Goals",
  yaxt = "n",
  xaxt = "n"
)

axis(1, at = 0:8)

axis(
  2,
  at = 1:4,
  labels = a2.plan$Measure,
  las = 1
)

segments(
  x0 = 0,
  y0 = 1:4,
  x1 = a2.plan$Months.To.A2,
  y1 = 1:4,
  lty = 2
)

Accelerated Progress to Reach A2 Benchmarks in Three Months

accelerated.plan <- a2.plan

accelerated.plan$New.Monthly.Goal <- ceiling(
  (accelerated.plan$A2.Target - accelerated.plan$Current) / 3
)

accelerated.plan$Increase.Needed <-
  accelerated.plan$New.Monthly.Goal - accelerated.plan$Monthly.Goal

accelerated.plan
percent.change <- (
  accelerated.plan$Increase.Needed /
  accelerated.plan$Monthly.Goal
) * 100

barplot(
  percent.change,
  names.arg = accelerated.plan$Measure,
  main = "Monthly Study Goal Changes for Three-Month Plan",
  xlab = "Language Learning Measure",
  ylab = "Required Change in Monthly Goal (%)",
  ylim = c(-100, 150),
  yaxt = "n"
)

axis(2, at = seq(-100, 150, 50))
abline(h = 0, lty = 2)

These results show that the current vocabulary and reaading goals are on trajectory to reach A2 goals in three months, however, listening and speaking hours need to increase 120% and 80% to reach A2 goals. With this data, the user can maintain reading and vocabulary but increase their listening and speaking goals. This shows us how R can identify weaknesses in language learning to adjust study plans based on goals that are measurable.

Further Resources

R Programming and Data Analysis

Google’s R Style Guide

Google’s R Style Guide

This guide provides recommendations for writing clear and consistent R code, including naming variables, spacing, and formatting. Following these conventions makes code easier to read and understand.

R Markdown Formatting and Presentation

Pimp My RMD

Pimp My RMD

This resource demonstrates ways to improve the appearance of R Markdown documents through formatting, themes, and layout options. It is useful for creating reports that are visually organized and easier for readers to navigate.

Syntax Highlighting Style in R Markdown

Syntax Highlighting Style in R Markdown

This resource explains how syntax highlighting can improve the readability of R code in R Markdown documents. Different highlighting styles make it easier for readers to distinguish functions, variables, and other elements of code.

Language Proficiency and CEFR

Common European Framework of Reference for Languages (CEFR)

Council of Europe – CEFR

The Council of Europe provides the CEFR framework for describing language proficiency from A1 through C2. This resource helps explain how language proficiency is evaluated across different skills. The vocabulary and activity targets used in this code-through are illustrative estimates rather than official CEFR requirements.

Works Cited

Holtz, Y. (2018, December 10). Pimp my RMD: A few tips for R Markdown. https://holtzy.github.io/Pimp-my-rmd/

Raviv, E. (2015, March 14). Syntax highlighting style in Rmarkdown. https://eranraviv.com/syntax-highlighting-style-in-rmarkdown/

Council of Europe. (2020). Common European Framework of Reference for Languages: Learning, teaching, assessment – Companion volume. Council of Europe Publishing. https://rm.coe.int/cefr-companion-volume-with-new-descriptors-2020/16809ea0d4