The goal of this capstone is to build an app that predicts the next word a user is about to type, similar to the keyboards on our phones. This report covers the first step: loading the English text data provided by SwiftKey (blogs, news articles and tweets), summarising it, and exploring which words and word combinations are most common. It ends with my plan for the prediction algorithm and the Shiny app.
url <- "https://d396qusza40orc.cloudfront.net/dsscapstone/dataset/Coursera-SwiftKey.zip"
zipfile <- "Coursera-SwiftKey.zip"
if (!file.exists(zipfile)) download.file(url, zipfile, mode = "wb")
if (!dir.exists("final")) unzip(zipfile)
read_text <- function(path) {
con <- file(path, open = "rb")
on.exit(close(con))
readLines(con, encoding = "UTF-8", skipNul = TRUE)
}
data_dir <- file.path("final", "en_US")
files <- file.path(data_dir, c("en_US.blogs.txt", "en_US.news.txt", "en_US.twitter.txt"))
texts <- list(
Blogs = read_text(files[1]),
News = read_text(files[2]),
Twitter = read_text(files[3])
)
words_per_line <- lapply(texts, stri_count_words)
summary_tbl <- data.frame(
Source = names(texts),
`Size (MB)` = round(file.info(files)$size / 1024^2, 1),
Lines = sapply(texts, length),
Words = sapply(words_per_line, sum),
`Words per line` = sapply(words_per_line, function(x) round(mean(x), 1)),
`Longest line (chars)` = sapply(texts, function(x) max(stri_length(x))),
check.names = FALSE
)
kable(summary_tbl, row.names = FALSE, format.args = list(big.mark = ","))
| Source | Size (MB) | Lines | Words | Words per line | Longest line (chars) |
|---|---|---|---|---|---|
| Blogs | 200.4 | 899,288 | 37,546,806 | 41.8 | 40,833 |
| News | 196.3 | 1,010,242 | 34,762,658 | 34.4 | 11,384 |
| 159.4 | 2,360,148 | 30,096,690 | 12.8 | 140 |
The three sources contain a similar number of words, but they are very different in style: tweets are short, while blog posts and news articles are much longer.
set.seed(2026)
sample_rate <- 0.02
sample_lines <- function(x) x[as.logical(rbinom(length(x), 1, sample_rate))]
clean_text <- function(x) {
x <- gsub("\u2019", "'", x) # curly apostrophe -> straight
x <- iconv(x, "UTF-8", "ASCII", sub = " ") # drop non-English characters
x <- tolower(x)
x <- gsub("(https?://|www\\.)\\S+", " ", x) # URLs
x <- gsub("\\S+@\\S+", " ", x) # e-mails
x <- gsub("[^a-z' ]", " ", x) # numbers and punctuation
x <- gsub("\\s+", " ", trimws(x))
x
}
corpus <- bind_rows(lapply(names(texts), function(s)
tibble(source = s, text = clean_text(sample_lines(texts[[s]])))
))
rm(texts); invisible(gc()) # free memory
The analysis below uses a random 2% sample of each source (85,058 lines in total), which is enough to see the main patterns while keeping the computation fast.
count_ngrams <- function(n) {
corpus %>%
unnest_tokens(ngram, text, token = "ngrams", n = n) %>%
filter(!is.na(ngram)) %>%
count(ngram, sort = TRUE)
}
unigrams <- count_ngrams(1)
bigrams <- count_ngrams(2)
trigrams <- count_ngrams(3)
plot_top <- function(df, title, k = 20, colour = "steelblue") {
df %>%
slice_head(n = k) %>%
ggplot(aes(x = reorder(ngram, n), y = n)) +
geom_col(fill = colour) +
coord_flip() +
labs(title = title, x = NULL, y = "Frequency in sample") +
theme_minimal()
}
plot_top(unigrams, "Top 20 words")
plot_top(bigrams, "Top 20 two-word phrases", colour = "darkorange")
plot_top(trigrams, "Top 20 three-word phrases", colour = "seagreen")
The most common words are short “function” words such as the, to and and. In many text analyses these are removed, but here they are kept on purpose: they are exactly the words people type most often, so the app should be able to predict them.
coverage <- unigrams %>%
mutate(rank = row_number(), cum_share = cumsum(n) / sum(n))
words_50 <- which(coverage$cum_share >= 0.5)[1]
words_90 <- which(coverage$cum_share >= 0.9)[1]
ggplot(coverage, aes(rank, cum_share)) +
geom_line(colour = "steelblue", linewidth = 1) +
geom_hline(yintercept = c(0.5, 0.9), linetype = "dashed", colour = "grey50") +
scale_x_log10(labels = scales::comma) +
scale_y_continuous(labels = scales::percent) +
labs(title = "Share of all text covered by the most frequent words",
x = "Number of distinct words (log scale)", y = "Coverage") +
theme_minimal()
Out of 73,496 distinct words in the sample, only 144 words cover 50% of all the text, and 7,112 cover 90%. A small vocabulary therefore goes a long way, which means the prediction model can be kept small and fast.