1. Summary for the busy reader

The goal of this capstone is a Shiny app that predicts the next word as a user types, much like a phone keyboard. This report shows that the data has been downloaded and loaded, summarizes the three text files, and outlines the plan for the prediction algorithm.

Main findings so far

  • The data is large: three English text files (blogs, news, Twitter) with millions of lines in total.
  • The sources differ in style. Tweets are short, news and blog lines are longer.
  • A small number of common words account for most of what people write, so a compact model can still be useful.
  • Word pairs and triples (n-grams) repeat often, which makes them a good basis for predicting the next word.

2. Loading the data

The data comes from the Coursera/SwiftKey corpus. Only the English (en_US) files are used.

dir.create("data", showWarnings = FALSE)
options(timeout = 7200)
zip_path <- "data/Coursera-SwiftKey.zip"
url <- "https://d396qusza40orc.cloudfront.net/dsscapstone/dataset/Coursera-SwiftKey.zip"

if (!file.exists("data/final/en_US/en_US.blogs.txt")) {
    if (!file.exists(zip_path)) download.file(url, zip_path, mode = "wb", method = "libcurl")
  unzip(zip_path, exdir = "data")
}

read_text <- function(f) {
  con <- file(f, "rb")
  on.exit(close(con))
  readLines(con, encoding = "UTF-8", skipNul = TRUE, warn = FALSE)
}

blogs   <- read_text("data/final/en_US/en_US.blogs.txt")
news    <- read_text("data/final/en_US/en_US.news.txt")
twitter <- read_text("data/final/en_US/en_US.twitter.txt")

3. Basic summary of the three files

summarize_source <- function(x, name, path) {
  words <- stri_count_words(x)
  data.frame(
    Source = name,
    `File size (MB)` = round(file.size(path) / 1024^2, 1),
    Lines = length(x),
    `Total words` = sum(words, na.rm = TRUE),
    `Mean words per line` = round(mean(words, na.rm = TRUE), 1),
    `Longest line (characters)` = max(nchar(x)),
    check.names = FALSE
  )
}

summary_tbl <- bind_rows(
  summarize_source(blogs,   "Blogs",   "data/final/en_US/en_US.blogs.txt"),
  summarize_source(news,    "News",    "data/final/en_US/en_US.news.txt"),
  summarize_source(twitter, "Twitter", "data/final/en_US/en_US.twitter.txt")
)
kable(summary_tbl, format.args = list(big.mark = ","),
      caption = "Line and word counts for each file")
Line and word counts for each file
Source File size (MB) Lines Total words Mean words per line Longest line (characters)
Blogs 200.4 899,288 37,546,806 41.8 40,833
News 196.3 1,010,242 34,762,658 34.4 11,384
Twitter 159.4 2,360,148 30,096,690 12.8 140

Because the files are big, the rest of the analysis uses a random 2% sample of each file. This is fast and still representative.

set.seed(123)
take <- function(x, p = 0.02) x[sample(length(x), round(length(x) * p))]

samp <- bind_rows(
  data.frame(source = "Blogs",   text = take(blogs),   stringsAsFactors = FALSE),
  data.frame(source = "News",    text = take(news),    stringsAsFactors = FALSE),
  data.frame(source = "Twitter", text = take(twitter), stringsAsFactors = FALSE)
)

# Basic cleaning: lower case, drop URLs, keep letters and apostrophes only
samp$text <- tolower(samp$text)
samp$text <- gsub("http\\S+|www\\.\\S+", " ", samp$text)
samp$text <- gsub("[^a-z' ]", " ", samp$text)
samp$text <- gsub("\\s+", " ", trimws(samp$text))
samp <- samp[nchar(samp$text) > 0, ]

4. How long are the lines?

samp$n_words <- stri_count_words(samp$text)

ggplot(samp, aes(n_words, fill = source)) +
  geom_histogram(binwidth = 2, show.legend = FALSE) +
  coord_cartesian(xlim = c(0, 120)) +
  facet_wrap(~source, scales = "free_y") +
  labs(title = "Words per line (2% sample)", x = "Words in a line", y = "Number of lines")

Takeaway: Tweets are short (capped by Twitter’s length limit), while blog lines can be very long. The app should work well on short, casual text as well as longer prose.

5. Which words are most common?

unigrams <- samp %>%
  unnest_tokens(word, text, token = "words") %>%
  count(word, sort = TRUE)

unigrams %>% slice_head(n = 15) %>%
  ggplot(aes(reorder(word, n), n)) +
  geom_col(fill = "steelblue") + coord_flip() +
  labs(title = "15 most frequent words", x = NULL, y = "Count")

How many words cover most of the text?

unigrams <- unigrams %>% mutate(cum_share = cumsum(n) / sum(n), rank = row_number())
w50 <- min(which(unigrams$cum_share >= 0.5))
w90 <- min(which(unigrams$cum_share >= 0.9))

ggplot(unigrams, aes(rank, cum_share)) +
  geom_line(color = "darkred", linewidth = 1) +
  scale_x_log10(labels = scales::comma) +
  scale_y_continuous(labels = scales::percent) +
  geom_hline(yintercept = c(0.5, 0.9), linetype = "dashed") +
  labs(title = "Share of all words covered by the most frequent words",
       x = "Number of unique words (log scale)", y = "Coverage of all word occurrences")

Out of 73,072 distinct words in the sample, only 143 words cover 50% of everything written and 7,053 words cover 90%. The rare words (the long tail) can be trimmed to keep the app small and fast.

6. Common word pairs and triples

top_ngrams <- function(n) {
  samp %>%
    unnest_tokens(ngram, text, token = "ngrams", n = n) %>%
    filter(!is.na(ngram)) %>%
    count(ngram, sort = TRUE)
}
bigrams  <- top_ngrams(2)
trigrams <- top_ngrams(3)

plot_top <- function(df, ttl) {
  df %>% slice_head(n = 15) %>%
    ggplot(aes(reorder(ngram, n), n)) +
    geom_col(fill = "seagreen") + coord_flip() +
    labs(title = ttl, x = NULL, y = "Count")
}
plot_top(bigrams, "15 most frequent two-word phrases")

plot_top(trigrams, "15 most frequent three-word phrases")

Takeaway: Everyday phrases such as “of the” and “thanks for the” repeat constantly. If a user has typed “thanks for”, the model can learn that “the” is a very likely next word.

7. Other things noticed

  • Mixed content: The data contains URLs, numbers, hashtags, emoji and some non-English text. These were removed or simplified during cleaning.
  • Profanity: Some text includes offensive words. The final app should filter these from its suggestions.
  • Unseen phrases: Many longer phrases appear only once. The model needs a fallback for word combinations it has never seen.

8. Plan for the prediction algorithm and Shiny app

Algorithm

  1. Build n-gram tables (1 to 4 words) from a larger, cleaned sample of all three sources.
  2. Remove very rare n-grams to keep the model small and fast.
  3. Predict the next word using stupid backoff (or Kneser-Ney smoothing): look for the longest matching phrase first, then back off to shorter ones.
  4. Evaluate accuracy on a held-out test set (top-1 and top-3 accuracy) and track speed and memory use.

Shiny app

  1. A text box where the user types a phrase.
  2. The top three predicted next words are shown instantly as clickable suggestions.
  3. A short help panel explains how to use the app.
  4. The app loads precomputed tables so that it responds quickly and fits within hosting limits.

Feedback requested: Is a 2-4 word context the right trade-off between accuracy and app size, and are there other features worth adding?