This report presents a analysis of the SwiftKey Dataset provided by Johns Hopkins and Coursera.
| Source | Lines | Size_MB |
|---|---|---|
| Blogs | 899288 | 200.4 |
| News | 1010242 | 196.3 |
| 2360148 | 159.4 |
The dataset is large — over 4 million lines and 556 MB of raw text. To keep analysis efficient, a random sample of 10,000 lines per source was used throughout this report.
Blogs contain the longest text averaging around 42 words per line, while Twitter is shortest at around 12 words due to character limits. This difference will influence how the prediction model handles short vs long context.
The most frequent words — time, people, day, love, life — reflect the personal and conversational nature of the data sources.
Trigrams like “one of the” and “a lot of” reveal natural English phrasing patterns that the prediction model will learn from.
The next word prediction model will work like:
The Shiny app will:
The analysis confirms the data is rich, diverse, and suitable for building a next word prediction. findings include significant differences in text length across sources, and clear patterns in word and phrase frequency that will directly inform the prediction algorithm.