Mobile device users generate billions of keystrokes daily across messaging, social media, and web browsing. Smart keyboard technology—powered by natural language processing (NLP) and predictive text models—streamlines mobile typing by accurately forecasting upcoming words.
This milestone report presents an exploratory data analysis (EDA) of the Coursera-SwiftKey English text corpus, which combines entries from blogs, news articles, and Twitter. We summarize key linguistic structures, word distributions, and n-gram frequencies, and establish an engineering blueprint for our final predictive text application.
The English dataset consists of three text files. Below is a summary table detailing the file sizes, line counts, character counts, and total word counts across the three corpora:
| Source | Size_MB | Lines | Words | Chars | Max_Line_Length | Avg_Words_Per_Line |
|---|---|---|---|---|---|---|
| blogs | 200.42 | 899288 | 37334131 | 206824505 | 40833 | 41.52 |
| news | 196.28 | 1010242 | 34372530 | 203223159 | 11384 | 34.02 |
| 159.36 | 2360148 | 30373583 | 162096241 | 140 | 12.87 |
Due to memory constraints and the need for fast iterative modeling, we draw a representative random sample (3%) from each corpus.
@username), numbers, and excess
whitespace.The distribution of single words exhibits classic Zipf’s Law behavior, where stopwords (e.g., the, to, and, of) dominate the total token count:
Analyzing pairs of consecutive words reveals high-frequency prepositional and determiner combinations:
Trigrams capture contextual conversational idioms and multi-word transitions:
A critical question in predictive text modeling is dictionary efficiency: how many unique words are required to cover 50% and 90% of all word occurrences in the corpus?
The exploratory analysis confirms that natural language follows highly skewed frequency distributions suitable for n-gram modeling. The next milestone will finalize the backoff prediction algorithm, test benchmark phrases, and deploy the interactive Shiny data product.