This report is a quick progress update on my Capstone project. I explored three English text datasets — blogs, news, and Twitter — and I’m sharing what I found so far, plus my early plan for a prediction algorithm and a Shiny app. The short version: word frequencies are heavily concentrated. About 107 common words cover 50% of all word occurrences. I’m planning to build an n-gram model and turn it into a Shiny app.
Here’s the basic summary for each dataset:
| Dataset | FileSize_MB | Lines | Words | AvgWordsPerLine |
|---|---|---|---|---|
| Blogs | 200.42 | 899288 | 37546806 | 41.75 |
| News | 196.28 | 1010206 | 34761151 | 34.41 |
| 159.36 | 2360148 | 30096690 | 12.75 |
| word | freq |
|---|---|
| the | 205748 |
| and | 120901 |
| to | 118716 |
| a | 99359 |
| of | 96920 |
| i | 85657 |
| in | 66229 |
| that | 50908 |
| is | 47821 |
| it | 44268 |
| for | 40104 |
| you | 32698 |
| with | 32027 |
| was | 30726 |
| on | 30612 |
| my | 29872 |
| this | 28765 |
| as | 24623 |
| have | 24180 |
| be | 23157 |
| word | freq |
|---|---|
| the | 194310 |
| to | 89274 |
| and | 87641 |
| a | 86740 |
| of | 76057 |
| in | 66698 |
| for | 34996 |
| that | 34082 |
| is | 28020 |
| on | 26318 |
| with | 25081 |
| said | 24889 |
| he | 22669 |
| was | 22609 |
| it | 21628 |
| at | 21105 |
| as | 18328 |
| i | 15695 |
| his | 15350 |
| be | 15120 |
| word | freq |
|---|---|
| the | 39856 |
| to | 33116 |
| i | 30247 |
| a | 25820 |
| you | 23210 |
| and | 18352 |
| for | 16266 |
| in | 16069 |
| of | 15161 |
| is | 15148 |
| it | 12235 |
| my | 12147 |
| on | 11716 |
| that | 10016 |
| me | 8300 |
| be | 7943 |
| at | 7800 |
| your | 7230 |
| with | 7221 |
| have | 6976 |
To cover 50% of all word occurrences, I need about 107 words. To cover 90%, I need about 6801 words.
I’m going to build a simple n-gram model. When a user types some text, the model will look at the previous 1–3 words and predict the most likely next word. If a particular word combination never showed up in the data, the model will fall back to shorter combinations.
I’ll build a Shiny app where users can type text in an input box and see the predicted next word in real time. The app will be deployed on shinyapps.io.