Twitter has the most lines (2.3M), blogs is largest at 200MB
## Warning in readLines(con, 10000): line 7155 appears to contain an embedded nul
## Warning in readLines(con, 10000): line 8547 appears to contain an embedded nul
## Warning in readLines(con, 10000): line 4086 appears to contain an embedded nul
## Warning in readLines(con, 10000): line 9032 appears to contain an embedded nul
## File Size_MB Lines
## en_US.blogs.txt en_US.blogs.txt 200.42 899288
## en_US.news.txt en_US.news.txt 196.28 1010242
## en_US.twitter.txt en_US.twitter.txt 159.36 2360148
5% probabilistic sampling with fixed seed for reproducibility.
## Blogs lines: 44917
## News lines: 50497
## Twitter lines: 118272
Twitter has most lines but fewer tokens — tweets are shorter.
## Blogs tokens: 1750622
## News tokens: 1614933
## Twitter tokens: 1391912
The appears 3-6% of all tokens; 61% of Twitter vocab appears once.
## --- BLOGS ---
## Total tokens: 1750622
## Unique words: 71837
##
## Top 20 words:
## Rank Word Freq Pct(%)
## 1 the 92861 5.304
## 2 and 54107 3.091
## 3 to 53016 3.028
## 4 of 44003 2.514
## 5 in 29678 1.695
## 6 that 23071 1.318
## 7 is 21568 1.232
## 8 it 19982 1.141
## 9 for 18133 1.036
## 10 you 14331 0.819
## 11 with 14110 0.806
## 12 was 13836 0.790
## 13 on 13775 0.787
## 14 my 13283 0.759
## 15 this 12853 0.734
## 16 as 11327 0.647
## 17 have 10964 0.626
## 18 be 10427 0.596
## 19 but 10163 0.581
## 20 are 9826 0.561
##
## Hapax (1x): 37064 - 51.59 % of vocab
## Dis (2x): 9398 - 13.08 % of vocab
## Tris (3x): 4793 - 6.67 % of vocab
## 10+ times: 10802
## --- NEWS ---
## Total tokens: 1614933
## Unique words: 70677
##
## Top 20 words:
## Rank Word Freq Pct(%)
## 1 the 98334 6.089
## 2 to 45210 2.799
## 3 and 43979 2.723
## 4 of 38314 2.372
## 5 in 33707 2.087
## 6 for 17485 1.083
## 7 that 17308 1.072
## 8 is 14066 0.871
## 9 on 13015 0.806
## 10 said 12700 0.786
## 11 with 12461 0.772
## 12 he 11561 0.716
## 13 was 11522 0.713
## 14 it 10908 0.675
## 15 at 10377 0.643
## 16 as 9131 0.565
## 17 his 7787 0.482
## 18 be 7630 0.472
## 19 from 7609 0.471
## 20 but 7519 0.466
##
## Hapax (1x): 34292 - 48.52 % of vocab
## Dis (2x): 9789 - 13.85 % of vocab
## Tris (3x): 5054 - 7.15 % of vocab
## 10+ times: 11526
## --- TWITTER ---
## Total tokens: 1391912
## Unique words: 69303
##
## Top 20 words:
## Rank Word Freq Pct(%)
## 1 the 46953 3.373
## 2 to 39253 2.820
## 3 you 27234 1.957
## 4 and 21774 1.564
## 5 for 19323 1.388
## 6 in 19019 1.366
## 7 of 18028 1.295
## 8 is 17794 1.278
## 9 it 14679 1.055
## 10 my 14413 1.035
## 11 on 13833 0.994
## 12 that 11666 0.838
## 13 me 10131 0.728
## 14 be 9256 0.665
## 15 at 9042 0.650
## 16 with 8650 0.621
## 17 your 8634 0.620
## 18 have 8385 0.602
## 19 so 8163 0.586
## 20 this 8058 0.579
##
## Hapax (1x): 42287 - 61.02 % of vocab
## Dis (2x): 7760 - 11.2 % of vocab
## Tris (3x): 3683 - 5.31 % of vocab
## 10+ times: 8212
Most frequent are function word combos like ‘of the’, ’to the.
## Testando com Twitter primeiro...
## --- TWITTER ---
## Unique bigrams (sample): 26505
## Top 10 bigrams:
## 1. to the 32
## 2. the the 30
## 3. the you 27
## 4. for the 24
## 5. the to 23
## 6. you the 21
## 7. to to 19
## 8. to and 18
## 9. to in 16
## 10. in to 15
##
## Unique trigrams (sample): 19973
## Top 10 trigrams:
## 1. for on be 3
## 2. all get to 2
## 3. by to to 2
## 4. dont what my 2
## 5. for of and 2
## 6. it you you 2
## 7. of and at 2
## 8. on the you 2
## 9. rt the be 2
## 10. the and the 2
## --- BLOGS ---
## Unique bigrams (sample): 25648
## Top 10 bigrams:
## 1. the the 99
## 2. the to 51
## 3. of the 50
## 4. the and 50
## 5. and the 48
## 6. the of 42
## 7. to the 41
## 8. the in 37
## 9. in the 33
## 10. to to 33
##
## Unique trigrams (sample): 19907
## Top 10 trigrams:
## 1. the the the 5
## 2. in the to 4
## 3. the it to 4
## 4. and the the 3
## 5. and this the 3
## 6. it to the 3
## 7. the in the 3
## 8. the of to 3
## 9. the the and 3
## 10. the the in 3
## --- NEWS ---
## Unique bigrams (sample): 26387
## Top 10 bigrams:
## 1. the the 103
## 2. of the 55
## 3. to the 54
## 4. and the 50
## 5. the to 43
## 6. the and 41
## 7. the in 41
## 8. the of 36
## 9. to and 28
## 10. of and 27
##
## Unique trigrams (sample): 19937
## Top 10 trigrams:
## 1. the the the 4
## 2. to the the 4
## 3. in to the 3
## 4. of and the 3
## 5. the are the 3
## 6. the to the 3
## 7. to the and 3
## 8. to to the 3
## 9. an this the 2
## 10. and of the 2
Just 124 words cover 50% of blogs; ~7,000 cover 90%.
## --- BLOGS ---
## Words for 50% coverage: 124 - word: right
## Words for 90% coverage: 7062 - word: ticking
##
## Progressive coverage:
## 10% -> 3 words
## 20% -> 9 words
## 30% -> 24 words
## 40% -> 56 words
## 50% -> 124 words
## 60% -> 313 words
## 70% -> 824 words
## 80% -> 2183 words
## 90% -> 7062 words
## 100% -> 71837 words
## --- NEWS ---
## Words for 50% coverage: 220 - word: health
## Words for 90% coverage: 8580 - word: stephens
##
## Progressive coverage:
## 10% -> 3 words
## 20% -> 10 words
## 30% -> 29 words
## 40% -> 78 words
## 50% -> 220 words
## 60% -> 561 words
## 70% -> 1311 words
## 80% -> 3026 words
## 90% -> 8580 words
## 100% -> 70677 words
## --- TWITTER ---
## Words for 50% coverage: 138 - word: down
## Words for 90% coverage: 5894 - word: bennett
##
## Progressive coverage:
## 10% -> 5 words
## 20% -> 14 words
## 30% -> 34 words
## 40% -> 68 words
## 50% -> 138 words
## 60% -> 290 words
## 70% -> 665 words
## 80% -> 1748 words
## 90% -> 5894 words
## 100% -> 69303 words
49K-53K rare words potentially non-English per dataset.
## --- BLOGS ---
## Total unique words: 71837
## Rare words (freq <= 3): 51255 ( 71.35 % of vocab)
## Estimated foreign/rare: 51255
## Sample:
## mcbrides, impresses, talentsfor, timeswe, barcis, filings, nudist, jennis, margo, spotlightsprouting, uptrend, punctuality
## --- NEWS ---
## Total unique words: 70677
## Rare words (freq <= 3): 49135 ( 69.52 % of vocab)
## Estimated foreign/rare: 49135
## Sample:
## minddestroying, harriman, twittercomrayrinaldi, vickerson, bahr, empties, outspsortcom, hosed, merin, terese, swansons, reynoso
## --- TWITTER ---
## Total unique words: 69303
## Rare words (freq <= 3): 53730 ( 77.53 % of vocab)
## Estimated foreign/rare: 53730
## Sample:
## metzger, pleaz, summermusic, tempter, canitbemondayalready, magicians, nunthing, queenstown, mebeen, yella, sonw, aut
Removing hapax cuts 50% vocab, loses only 2-3% of tokens.
## 1 - Remove hapax (words appearing once):
## BLOGS: 37064 hapax (51.6% vocab), 2.12% of tokens
## NEWS: 34292 hapax (48.5% vocab), 2.12% of tokens
## TWITTER: 42287 hapax (61.0% vocab), 3.04% of tokens
##
## 2 - Stemming: group 'running/runs/ran' -> 'run' (~30-40% vocab reduction)
##
## 3 - Reduced dictionary:
## BLOGS: 7062=90%, 16521=95%, 54331=99% coverage
## NEWS: 8580=90%, 18614=95%, 54528=99% coverage
## TWITTER: 5894=90%, 15389=95%, 55384=99% coverage
##
## 4 - Normalize numbers -> <NUM>, URLs -> <URL>
Contexto para os 3 gráficos (barras, cobertura, Zipf).
~70K unique words per dataset; 6K-8.5K covers 90%.
## Metric Blogs News Twitter
## Sample lines 44917.00 50497.00 118272.00
## Sample tokens 1750622.00 1614933.00 1391912.00
## Unique words 71837.00 70677.00 69303.00
## Hapax (1x) 37064.00 34292.00 42287.00
## Hapax % of vocab 51.59 48.52 61.02
## Words for 50% cov. 124.00 220.00 138.00
## Words for 90% cov. 7062.00 8580.00 5894.00
## Est. foreign words 51255.00 49135.00 53730.00
Based on the findings from this exploratory analysis, I have defined a clear plan for building the prediction algorithm and the interactive application.
I will build a 4-gram model — meaning it will look at up to three previous words to predict the next one. The model will use Stupid Backoff scoring, an efficient technique for handling word combinations that were not seen in the training data. For rare n-grams, Katz backoff with discounting will smooth the probability estimates. The model will be stored using data.table objects for fast lookups. To keep the model under 100MB of RAM — essential for running on Shiny’s servers — I will apply stemming and remove words that appear only once, which reduces vocabulary by over 50% while losing only 2-3% of token coverage. Profanity and extremely rare words will also be filtered out.
The final product will be a simple, fast web application where:
To measure how well the model performs, I will: