Overview

The goal of this capstone is to build an app that predicts the next word a user is about to type, similar to the keyboards on our phones. This report covers the first step: loading the English text data provided by SwiftKey (blogs, news articles and tweets), summarising it, and exploring which words and word combinations are most common. It ends with my plan for the prediction algorithm and the Shiny app.

1. Loading the data

url     <- "https://d396qusza40orc.cloudfront.net/dsscapstone/dataset/Coursera-SwiftKey.zip"
zipfile <- "Coursera-SwiftKey.zip"
if (!file.exists(zipfile)) download.file(url, zipfile, mode = "wb")
if (!dir.exists("final")) unzip(zipfile)

read_text <- function(path) {
  con <- file(path, open = "rb")
  on.exit(close(con))
  readLines(con, encoding = "UTF-8", skipNul = TRUE)
}

data_dir <- file.path("final", "en_US")
files <- file.path(data_dir, c("en_US.blogs.txt", "en_US.news.txt", "en_US.twitter.txt"))

texts <- list(
  Blogs   = read_text(files[1]),
  News    = read_text(files[2]),
  Twitter = read_text(files[3])
)

2. Basic summary

words_per_line <- lapply(texts, stri_count_words)

summary_tbl <- data.frame(
  Source           = names(texts),
  `Size (MB)`      = round(file.info(files)$size / 1024^2, 1),
  Lines            = sapply(texts, length),
  Words            = sapply(words_per_line, sum),
  `Words per line` = sapply(words_per_line, function(x) round(mean(x), 1)),
  `Longest line (chars)` = sapply(texts, function(x) max(stri_length(x))),
  check.names = FALSE
)

kable(summary_tbl, row.names = FALSE, format.args = list(big.mark = ","))
Source Size (MB) Lines Words Words per line Longest line (chars)
Blogs 200.4 899,288 37,546,806 41.8 40,833
News 196.3 1,010,242 34,762,658 34.4 11,384
Twitter 159.4 2,360,148 30,096,690 12.8 140

The three sources contain a similar number of words, but they are very different in style: tweets are short, while blog posts and news articles are much longer.

3. Sampling and cleaning

set.seed(2026)
sample_rate <- 0.02

sample_lines <- function(x) x[as.logical(rbinom(length(x), 1, sample_rate))]

clean_text <- function(x) {
  x <- gsub("\u2019", "'", x)                      # curly apostrophe -> straight
  x <- iconv(x, "UTF-8", "ASCII", sub = " ")       # drop non-English characters
  x <- tolower(x)
  x <- gsub("(https?://|www\\.)\\S+", " ", x)      # URLs
  x <- gsub("\\S+@\\S+", " ", x)                   # e-mails
  x <- gsub("[^a-z' ]", " ", x)                    # numbers and punctuation
  x <- gsub("\\s+", " ", trimws(x))
  x
}

corpus <- bind_rows(lapply(names(texts), function(s)
  tibble(source = s, text = clean_text(sample_lines(texts[[s]])))
))

rm(texts); invisible(gc())   # free memory

The analysis below uses a random 2% sample of each source (85,058 lines in total), which is enough to see the main patterns while keeping the computation fast.

4. Most frequent words and phrases

count_ngrams <- function(n) {
  corpus %>%
    unnest_tokens(ngram, text, token = "ngrams", n = n) %>%
    filter(!is.na(ngram)) %>%
    count(ngram, sort = TRUE)
}

unigrams <- count_ngrams(1)
bigrams  <- count_ngrams(2)
trigrams <- count_ngrams(3)

plot_top <- function(df, title, k = 20, colour = "steelblue") {
  df %>%
    slice_head(n = k) %>%
    ggplot(aes(x = reorder(ngram, n), y = n)) +
    geom_col(fill = colour) +
    coord_flip() +
    labs(title = title, x = NULL, y = "Frequency in sample") +
    theme_minimal()
}
plot_top(unigrams, "Top 20 words")

plot_top(bigrams, "Top 20 two-word phrases", colour = "darkorange")

plot_top(trigrams, "Top 20 three-word phrases", colour = "seagreen")

The most common words are short “function” words such as the, to and and. In many text analyses these are removed, but here they are kept on purpose: they are exactly the words people type most often, so the app should be able to predict them.

5. How many words do we need?

coverage <- unigrams %>%
  mutate(rank = row_number(), cum_share = cumsum(n) / sum(n))

words_50 <- which(coverage$cum_share >= 0.5)[1]
words_90 <- which(coverage$cum_share >= 0.9)[1]

ggplot(coverage, aes(rank, cum_share)) +
  geom_line(colour = "steelblue", linewidth = 1) +
  geom_hline(yintercept = c(0.5, 0.9), linetype = "dashed", colour = "grey50") +
  scale_x_log10(labels = scales::comma) +
  scale_y_continuous(labels = scales::percent) +
  labs(title = "Share of all text covered by the most frequent words",
       x = "Number of distinct words (log scale)", y = "Coverage") +
  theme_minimal()

Out of 73,496 distinct words in the sample, only 144 words cover 50% of all the text, and 7,112 cover 90%. A small vocabulary therefore goes a long way, which means the prediction model can be kept small and fast.

6. Plan for the prediction app

  1. Build n-gram tables. Using a larger sample, I will count 1-, 2-, 3- and 4-word sequences and drop very rare ones to keep the tables small.
  2. Predict with a “back-off” model. Given the last three words the user typed, the app looks for 4-word phrases starting with them and suggests the most frequent next word. If none is found, it “backs off” to the last two words, then the last word, and finally suggests the most common words overall.
  3. Shiny app. A simple web page with a text box: as the user types, it shows the most likely next words.
  4. Check accuracy and speed on text not used to build the model, balancing accuracy against the size of the model so the app responds instantly.