This report summarizes exploratory analysis of the SwiftKey dataset and outlines the design of a predictive text algorithm using n-gram models. The goal is to create an efficient next-word predictor suitable for deployment in a Shiny web application.
# Load datasets (adjust paths if needed)
blogs <- readLines("en_US/en_US.blogs.txt", encoding = "UTF-8", skipNul = TRUE)
news <- readLines("en_US/en_US.news.txt", encoding = "UTF-8", skipNul = TRUE)
## Warning in readLines("en_US/en_US.news.txt", encoding = "UTF-8", skipNul =
## TRUE): incomplete final line found on 'en_US/en_US.news.txt'
twitter <- readLines("en_US/en_US.twitter.txt", encoding = "UTF-8", skipNul = TRUE)
# Basic statistics
data_summary <- data.frame(
Source = c("Blogs", "News", "Twitter"),
Lines = c(length(blogs), length(news), length(twitter)),
Words = c(sum(str_count(blogs, "\\w+")),
sum(str_count(news, "\\w+")),
sum(str_count(twitter, "\\w+")))
)
knitr::kable(data_summary)
| Source | Lines | Words |
|---|---|---|
| Blogs | 899288 | 38309620 |
| News | 77259 | 2741594 |
| 2360148 | 31003544 |
set.seed(123)
sample_data <- c(sample(blogs, 10000),
sample(news, 10000),
sample(twitter, 10000))
sample_clean <- tolower(sample_data)
sample_clean <- gsub("[^a-z\\s]", "", sample_clean)
text_df <- tibble(line = 1:length(sample_clean), text = sample_clean)
# Unigram
unigrams <- text_df %>%
unnest_tokens(word, text) %>%
count(word, sort = TRUE)
# Plot top 20 unigrams
unigrams %>% top_n(20) %>%
ggplot(aes(reorder(word, n), n)) +
geom_col(fill = "steelblue") +
coord_flip() +
labs(title = "Top 20 Unigrams", x = "Word", y = "Frequency")
## Selecting by n
# Bigrams
bigrams <- text_df %>%
unnest_tokens(bigram, text, token = "ngrams", n = 2) %>%
count(bigram, sort = TRUE) %>%
separate(bigram, into = c("w1", "w2"))
# Trigrams
trigrams <- text_df %>%
unnest_tokens(trigram, text, token = "ngrams", n = 3) %>%
count(trigram, sort = TRUE) %>%
separate(trigram, into = c("w1", "w2", "w3"))
predict_next_word <- function(input) {
input <- tolower(gsub("[^a-z\\s]", "", input))
words <- tail(unlist(strsplit(input, " ")), 2)
if (length(words) == 2) {
match <- trigrams %>%
filter(w1 == words[1], w2 == words[2]) %>%
arrange(desc(n)) %>%
slice(1)
if (nrow(match) > 0) return(match$w3)
}
if (length(words) >= 1) {
match <- bigrams %>%
filter(w1 == words[length(words)]) %>%
arrange(desc(n)) %>%
slice(1)
if (nrow(match) > 0) return(match$w2)
}
return(unigrams$word[1]) # Fallback to most common unigram
}
predict_next_word("i love")
## [1] "thankyou"
predict_next_word("the weather")
## [1] "thankyou"
predict_next_word("new york")
## [1] "thankyou"
Storage Efficiency: Store only top 10,000
n-grams per level. Use data.table and save as
.rds for fast access.
Smoothing: Add-one (Laplace) smoothing can be added to assign non-zero probabilities to unseen n-grams.
Backoff Model: Implement a backoff from trigram → bigram → unigram if no match is found.
Memory Tools:
object.size() to measure sizegc() for memory cleanupRprof() or profvis::profvis() for
profilingThe Shiny app will:
reactiveVal, memoise)This report demonstrates the creation of a basic predictive text model using n-grams and a backoff algorithm. The next steps are to improve accuracy with smoothing and test the model on held-out data to evaluate top-N accuracy and user-perceived performance.
```