1. Executive Summary

Mobile device users generate billions of keystrokes daily across messaging, social media, and web browsing. Smart keyboard technology—powered by natural language processing (NLP) and predictive text models—streamlines mobile typing by accurately forecasting upcoming words.

This milestone report presents an exploratory data analysis (EDA) of the Coursera-SwiftKey English text corpus, which combines entries from blogs, news articles, and Twitter. We summarize key linguistic structures, word distributions, and n-gram frequencies, and establish an engineering blueprint for our final predictive text application.


2. Data Acquisition & Summary Statistics

The English dataset consists of three text files. Below is a summary table detailing the file sizes, line counts, character counts, and total word counts across the three corpora:

Table 1: Overview and Summary Statistics of English Corpora
Source Size_MB Lines Words Chars Max_Line_Length Avg_Words_Per_Line
blogs 200.42 899288 37334131 206824505 40833 41.52
news 196.28 1010242 34372530 203223159 11384 34.02
twitter 159.36 2360148 30373583 162096241 140 12.87

Key Observations:

  1. Twitter: Contains the largest number of lines (~2.36M) but the shortest average line length (~12.8 words) due to character limits.
  2. Blogs: Demonstrates the longest individual paragraphs, with a maximum line exceeding 40,000 characters.
  3. News: Features formal syntax, consistent sentence structures, and a moderate average line length (~34.4 words).

3. Sampling & Text Preprocessing

Due to memory constraints and the need for fast iterative modeling, we draw a representative random sample (3%) from each corpus.

Preprocessing Pipeline:

  • Cleaning & Lowercasing: Normalized all text to standard lowercase ASCII.
  • Punctuation & Entities: Removed URLs, email addresses, Twitter handles (@username), numbers, and excess whitespace.
  • Profanity Filtering: Filtered out profane and offensive words so they are never recommended to users.

4. Exploratory Data Analysis & Visualizations

4.1 Most Frequent Unigrams (Single Words)

The distribution of single words exhibits classic Zipf’s Law behavior, where stopwords (e.g., the, to, and, of) dominate the total token count:

4.2 Most Frequent Bigrams (Word Pairs)

Analyzing pairs of consecutive words reveals high-frequency prepositional and determiner combinations:

4.3 Most Frequent Trigrams (Three-Word Phrases)

Trigrams capture contextual conversational idioms and multi-word transitions:


5. Vocabulary Coverage Analysis

A critical question in predictive text modeling is dictionary efficiency: how many unique words are required to cover 50% and 90% of all word occurrences in the corpus?

  • 50% of all word instances can be covered by only 131 words.
  • 90% coverage is achieved with approximately 5517 words.
  • This heavy concentration allows aggressive model pruning, cutting memory by >90% while keeping high hit rates.

6. Next Steps & Shiny App Architecture

Prediction Algorithm: Stupid Backoff N-Gram Model

  1. Quadgram Lookup: Match the previous 3 words typed.
  2. Backoff: If no match or low confidence, back off to 3-grams, then 2-grams (\(S = 0.4 \times S_{n-1}\)), and finally top unigrams.
  3. Pruning & Indexing: Pre-sort and retain only top 3–5 completions per prefix to keep the deployed model under 10 MB RAM with sub-10ms response latency.

Planned Shiny Web Application Interface:

  • Real-Time Input Box: Reactive text entry that detects keystrokes.
  • Top-3 Word Suggestions: Interactive buttons to click/tap to complete the word.
  • Execution Telemetry: Visual display showing whether the prediction originated from a 4-gram, 3-gram, 2-gram, or baseline unigram.

7. Conclusion

The exploratory analysis confirms that natural language follows highly skewed frequency distributions suitable for n-gram modeling. The next milestone will finalize the backoff prediction algorithm, test benchmark phrases, and deploy the interactive Shiny data product.