Learning a new language is not as simple as studying words, it
requires progress in multiple areas. Simply tracking the amount of time
that someone has studied does not accurately show their progress or
aptitude. It is important to track and identify vocabulary mastered,
words read, listening hours, speaking hours, and days studied.
Collecting and organizing this data can help identify current standing,
measure progress, and potentially forecast future proficiency.
This code through will demonstrate how R can be used to analyze
language learning data. As an example, we will use the Hungarian
language and track vocabulary mastery, words read, listening hours,
speaking hours and days studied. These measures will be used to identify
and visualize current status, progress over time, potential future
progress and areas that need additional attention.
Language learning is not a simple task, especially self directed
language learning. It is important to follow a plan and track progress
over time, however, it can be difficult to identify areas for
improvement and current progress without proper data. For example, you
might feel like you are reading and understanding a lot, but when you go
to have a conversation, you cannot think of the words to use or how to
use them. This can be defeating and halt progress. Tracking language
learning data can help organize progress, identify areas that need
improvement, and motivate a student by showing how much they have
progressed over time and what their potential future progress could look
like.
Specifically, you’ll learn how to:
In order to track language learning progress, we first need to create
a dataset that measures progress consistently over time. For this
example, we will track five areas of Hungarian language learning:
vocabulary mastered, words read, listening hours, speaking hours, and
days studied.
Language learning progress can be measured in multiple ways. For this example, six months of learning data will be used to demonstrate how progress can be tracked over time in a simple way. The dataset includes cumulative vocabulary mastered, words read, listening hours, and speaking hours, along with the number of days studied each month.
I compare my language learning activity data to estimated CEFR proficiency benchmarks to see whether all variables are consistent with the same general proficiency level. The CEFR ranges from A1, representing a beginner language user, through C2, representing a highly proficient language user. This comparison does not denote a CEFR level exactly, but it can show whether each area of language learning falls within the general range expected for that level.
First, we will make a dataset containing six months of language learning activity. Each row represents one month of study, while the columns represent the different measures that are being tracked.
language.progress <- data.frame(
month = c(1, 2, 3, 4, 5, 6),
vocabulary = c(300, 550, 800, 1050, 1300, 1600),
words.read = c(5000, 12000, 20000, 29000, 39000, 50835),
listening.hours = c(3, 6, 9, 12, 16, 20),
speaking.hours = c(0.5, 1, 1.5, 2, 2.5, 3),
days.studied = c(12, 14, 15, 16, 18, 15)
)
language.progressNow that we have simple data, we can see how language learning has changed over the six month time frame. The following graph shows growth in vocabulary over time.
plot(language.progress$month,
language.progress$vocabulary,
type = "o",
xlab = "Month",
ylab = "Vocabulary Mastered",
main = "Vocabulary Progress Over Six Months",
ylim = c(0, 1600),
yaxt = "n")
axis(2, at = c(0, 400, 800, 1200, 1600))
Now that we can visualize vocabulary over time, we can compare
vocabulary with estimated CEFR benchmarks. These ranges do not represent
an official CEFR assesment, but provide context for the data.
vocabulary.current <- 1600
if (vocabulary.current < 500) {
vocabulary.level <- "Below A1"
} else if (vocabulary.current < 1000) {
vocabulary.level <- "A1"
} else if (vocabulary.current < 2000) {
vocabulary.level <- "A2"
} else if (vocabulary.current < 5000) {
vocabulary.level <- "B1"
} else if (vocabulary.current < 8000) {
vocabulary.level <- "B2"
} else if (vocabulary.current < 15000) {
vocabulary.level <- "C1"
} else {
vocabulary.level <- "C2"
}
vocabulary.level## [1] "A2"
Another way to measure language learning is through number of words read. This identifies the cumulative number of Hungarian words read during the six month period.
Total number of words read can be compared with the estimated CEFR reading benchmarks. This will help us to understand if reading exposure is progressing at the same rate as vocabulary.
## [1] "A1"
Listening exposure measures time spent hearing the language, which can be tracked to determine if listening practice is progressing at the same rate as other variables. Listening hours can be compared with estimated CEFR benchmarks.
## [1] "Below A1"
Speaking hours represents active language use. Speaking hours can develop differently and much later than other language learning variables. This can identify if active language use is keepinig pace with the other areas of learning. Speaking hours can be compared with estimated CEFR benchmarks to identify if active language use is progressing at a similar rate to the other measures.
## [1] "Below A1"
Days of language learning study can help to provide context for the other measurements. This variable will not be measured against a CEFR benchmark. This will help identify consistency of practice during the month.
plot(language.progress$month,
language.progress$days.studied,
type = "o",
xlab = "Month",
ylab = "Days Studied",
main = "Study Consistency Over Six Months",
ylim = c(0, 20),
yaxt = "n")
axis(2, at = c(0, 5, 10, 15, 20))
Vocabulary, words read, listening hours, and speaking hours are measured differently and cannot be direclty compared on the same scale. Instead, we can convert each measure to a percentage of the estimated benchmark to compare how each area has progressed over time.
progress.percent <- data.frame(
Month = language.progress$month,
Vocabulary = (language.progress$vocabulary / 2000) * 100,
Words.Read = (language.progress$words.read / 100000) * 100,
Listening = (language.progress$listening.hours / 150) * 100,
Speaking = (language.progress$speaking.hours / 30) * 100
)
progress.percentmatplot(
progress.percent$Month,
progress.percent[, 2:5],
type = "o",
lty = 1,
pch = 1:4,
col = c("#4472C4", "#70AD47", "#ED7D31", "#8064A2"),
xlab = "Month",
ylab = "Percent of A2 Benchmark",
main = "Language Learning Progress Over Six Months",
ylim = c(0, 100),
yaxt = "n"
)
axis(
2,
at = c(0, 20, 40, 60, 80, 100),
labels = c("0%", "20%", "40%", "60%", "80%", "100%")
)
legend(
"topleft",
legend = c("Vocabulary", "Words Read", "Listening", "Speaking"),
lty = 1,
pch = 1:4,
col = c("#4472C4", "#70AD47", "#ED7D31", "#8064A2")
)
Visualizing multiple measures together can provide a clear picture toward an estimated overall proficiency. Vocabulary falls withing the A2 range, words read are closer to A1 and listening and speaking remain below A1. We may be able to describe the overall estimate to be around A1 without taking a certified CEFR assessment. This can help tailor the language learning plan prior to an assessment.
overall.progress <- data.frame(
Measure = c("Vocabulary", "Words Read", "Listening",
"Speaking", "Overall Estimate"),
Estimated.Level = c(vocabulary.level, reading.level,
listening.level, speaking.level,
"Around A1")
)
overall.progressThe above information can help to tailor the language learning plan. R can be used to estimate how long it may take to reach learning goals. Firt, we can compare current measurements with target values and monthly progress to calculate the approximate number of months needed for each variable to achieve A2 in six months. Then, we can identify how to alter the plan to reach A2 level across all variables in three months. These estimates do not guarantee achievement of official CEFR levels.
a2.plan <- data.frame(
Measure = c("Vocabulary", "Words Read", "Listening", "Speaking"),
Current = c(1600, 50835, 20, 3),
A2.Target = c(2000, 100000, 150, 30),
Monthly.Goal = c(300, 20000, 20, 5)
)
a2.plan$Months.To.A2 <- ceiling(
(a2.plan$A2.Target - a2.plan$Current) / a2.plan$Monthly.Goal
)
a2.planplot(
a2.plan$Months.To.A2,
1:4,
pch = 19,
cex = 1.5,
xlim = c(0, 8),
ylim = c(0.5, 4.5),
xlab = "Months From Now",
ylab = "",
main = "Projected Timeline Toward A2 Learning Goals",
yaxt = "n",
xaxt = "n"
)
axis(1, at = 0:8)
axis(
2,
at = 1:4,
labels = a2.plan$Measure,
las = 1
)
segments(
x0 = 0,
y0 = 1:4,
x1 = a2.plan$Months.To.A2,
y1 = 1:4,
lty = 2
)accelerated.plan <- a2.plan
accelerated.plan$New.Monthly.Goal <- ceiling(
(accelerated.plan$A2.Target - accelerated.plan$Current) / 3
)
accelerated.plan$Increase.Needed <-
accelerated.plan$New.Monthly.Goal - accelerated.plan$Monthly.Goal
accelerated.planpercent.change <- (
accelerated.plan$Increase.Needed /
accelerated.plan$Monthly.Goal
) * 100
barplot(
percent.change,
names.arg = accelerated.plan$Measure,
main = "Monthly Study Goal Changes for Three-Month Plan",
xlab = "Language Learning Measure",
ylab = "Required Change in Monthly Goal (%)",
ylim = c(-100, 150),
yaxt = "n"
)
axis(2, at = seq(-100, 150, 50))
abline(h = 0, lty = 2)These results show that the current vocabulary and reaading goals are on trajectory to reach A2 goals in three months, however, listening and speaking hours need to increase 120% and 80% to reach A2 goals. With this data, the user can maintain reading and vocabulary but increase their listening and speaking goals. This shows us how R can identify weaknesses in language learning to adjust study plans based on goals that are measurable.
Google’s R Style Guide
This guide provides recommendations for writing clear and consistent R code, including naming variables, spacing, and formatting. Following these conventions makes code easier to read and understand.
Pimp My RMD
This resource demonstrates ways to improve the appearance of R Markdown documents through formatting, themes, and layout options. It is useful for creating reports that are visually organized and easier for readers to navigate.
Syntax Highlighting Style in R Markdown
Syntax Highlighting Style in R Markdown
This resource explains how syntax highlighting can improve the readability of R code in R Markdown documents. Different highlighting styles make it easier for readers to distinguish functions, variables, and other elements of code.
Common European Framework of Reference for Languages (CEFR)
The Council of Europe provides the CEFR framework for describing
language proficiency from A1 through C2. This resource helps explain how
language proficiency is evaluated across different skills. The
vocabulary and activity targets used in this code-through are
illustrative estimates rather than official CEFR requirements.
Holtz, Y. (2018, December 10). Pimp my RMD: A few tips for R Markdown. https://holtzy.github.io/Pimp-my-rmd/
Raviv, E. (2015, March 14). Syntax highlighting style in Rmarkdown. https://eranraviv.com/syntax-highlighting-style-in-rmarkdown/
Council of Europe. (2020). Common European Framework of Reference for Languages: Learning, teaching, assessment – Companion volume. Council of Europe Publishing. https://rm.coe.int/cefr-companion-volume-with-new-descriptors-2020/16809ea0d4