1. Basic File Statistics

Twitter has the most lines (2.3M), blogs is largest at 200MB

## Warning in readLines(con, 10000): line 7155 appears to contain an embedded nul
## Warning in readLines(con, 10000): line 8547 appears to contain an embedded nul
## Warning in readLines(con, 10000): line 4086 appears to contain an embedded nul
## Warning in readLines(con, 10000): line 9032 appears to contain an embedded nul
##                                File Size_MB   Lines
## en_US.blogs.txt     en_US.blogs.txt  200.42  899288
## en_US.news.txt       en_US.news.txt  196.28 1010242
## en_US.twitter.txt en_US.twitter.txt  159.36 2360148

2. Load Representative Sample (5%)

5% probabilistic sampling with fixed seed for reproducibility.

## Blogs lines: 44917
## News lines: 50497
## Twitter lines: 118272

3. Tokenization

Twitter has most lines but fewer tokens — tweets are shorter.

## Blogs tokens: 1750622
## News tokens: 1614933
## Twitter tokens: 1391912

4. Q1: Word Frequency Distribution

The appears 3-6% of all tokens; 61% of Twitter vocab appears once.

## --- BLOGS ---
## Total tokens: 1750622 
## Unique words: 71837 
## 
## Top 20 words:
## Rank  Word                 Freq       Pct(%)
## 1     the                  92861      5.304
## 2     and                  54107      3.091
## 3     to                   53016      3.028
## 4     of                   44003      2.514
## 5     in                   29678      1.695
## 6     that                 23071      1.318
## 7     is                   21568      1.232
## 8     it                   19982      1.141
## 9     for                  18133      1.036
## 10    you                  14331      0.819
## 11    with                 14110      0.806
## 12    was                  13836      0.790
## 13    on                   13775      0.787
## 14    my                   13283      0.759
## 15    this                 12853      0.734
## 16    as                   11327      0.647
## 17    have                 10964      0.626
## 18    be                   10427      0.596
## 19    but                  10163      0.581
## 20    are                  9826       0.561
## 
## Hapax (1x): 37064 - 51.59 % of vocab
## Dis (2x): 9398 - 13.08 % of vocab
## Tris (3x): 4793 - 6.67 % of vocab
## 10+ times: 10802
## --- NEWS ---
## Total tokens: 1614933 
## Unique words: 70677 
## 
## Top 20 words:
## Rank  Word                 Freq       Pct(%)
## 1     the                  98334      6.089
## 2     to                   45210      2.799
## 3     and                  43979      2.723
## 4     of                   38314      2.372
## 5     in                   33707      2.087
## 6     for                  17485      1.083
## 7     that                 17308      1.072
## 8     is                   14066      0.871
## 9     on                   13015      0.806
## 10    said                 12700      0.786
## 11    with                 12461      0.772
## 12    he                   11561      0.716
## 13    was                  11522      0.713
## 14    it                   10908      0.675
## 15    at                   10377      0.643
## 16    as                   9131       0.565
## 17    his                  7787       0.482
## 18    be                   7630       0.472
## 19    from                 7609       0.471
## 20    but                  7519       0.466
## 
## Hapax (1x): 34292 - 48.52 % of vocab
## Dis (2x): 9789 - 13.85 % of vocab
## Tris (3x): 5054 - 7.15 % of vocab
## 10+ times: 11526
## --- TWITTER ---
## Total tokens: 1391912 
## Unique words: 69303 
## 
## Top 20 words:
## Rank  Word                 Freq       Pct(%)
## 1     the                  46953      3.373
## 2     to                   39253      2.820
## 3     you                  27234      1.957
## 4     and                  21774      1.564
## 5     for                  19323      1.388
## 6     in                   19019      1.366
## 7     of                   18028      1.295
## 8     is                   17794      1.278
## 9     it                   14679      1.055
## 10    my                   14413      1.035
## 11    on                   13833      0.994
## 12    that                 11666      0.838
## 13    me                   10131      0.728
## 14    be                   9256       0.665
## 15    at                   9042       0.650
## 16    with                 8650       0.621
## 17    your                 8634       0.620
## 18    have                 8385       0.602
## 19    so                   8163       0.586
## 20    this                 8058       0.579
## 
## Hapax (1x): 42287 - 61.02 % of vocab
## Dis (2x): 7760 - 11.2 % of vocab
## Tris (3x): 3683 - 5.31 % of vocab
## 10+ times: 8212

5. Q2: 2-Gram and 3-Gram Frequencies

Most frequent are function word combos like ‘of the’, ’to the.

## Testando com Twitter primeiro...
## --- TWITTER ---
## Unique bigrams (sample): 26505 
## Top 10 bigrams:
##  1. to the                         32
##  2. the the                        30
##  3. the you                        27
##  4. for the                        24
##  5. the to                         23
##  6. you the                        21
##  7. to to                          19
##  8. to and                         18
##  9. to in                          16
## 10. in to                          15
## 
## Unique trigrams (sample): 19973 
## Top 10 trigrams:
##  1. for on be                                3
##  2. all get to                               2
##  3. by to to                                 2
##  4. dont what my                             2
##  5. for of and                               2
##  6. it you you                               2
##  7. of and at                                2
##  8. on the you                               2
##  9. rt the be                                2
## 10. the and the                              2
## --- BLOGS ---
## Unique bigrams (sample): 25648 
## Top 10 bigrams:
##  1. the the                        99
##  2. the to                         51
##  3. of the                         50
##  4. the and                        50
##  5. and the                        48
##  6. the of                         42
##  7. to the                         41
##  8. the in                         37
##  9. in the                         33
## 10. to to                          33
## 
## Unique trigrams (sample): 19907 
## Top 10 trigrams:
##  1. the the the                              5
##  2. in the to                                4
##  3. the it to                                4
##  4. and the the                              3
##  5. and this the                             3
##  6. it to the                                3
##  7. the in the                               3
##  8. the of to                                3
##  9. the the and                              3
## 10. the the in                               3
## --- NEWS ---
## Unique bigrams (sample): 26387 
## Top 10 bigrams:
##  1. the the                        103
##  2. of the                         55
##  3. to the                         54
##  4. and the                        50
##  5. the to                         43
##  6. the and                        41
##  7. the in                         41
##  8. the of                         36
##  9. to and                         28
## 10. of and                         27
## 
## Unique trigrams (sample): 19937 
## Top 10 trigrams:
##  1. the the the                              4
##  2. to the the                               4
##  3. in to the                                3
##  4. of and the                               3
##  5. the are the                              3
##  6. the to the                               3
##  7. to the and                               3
##  8. to to the                                3
##  9. an this the                              2
## 10. and of the                               2

6. Q3: Dictionary Coverage (50% and 90%)

Just 124 words cover 50% of blogs; ~7,000 cover 90%.

## --- BLOGS ---
## Words for 50% coverage: 124 - word: right 
## Words for 90% coverage: 7062 - word: ticking 
## 
## Progressive coverage:
##    10% -> 3 words
##    20% -> 9 words
##    30% -> 24 words
##    40% -> 56 words
##    50% -> 124 words
##    60% -> 313 words
##    70% -> 824 words
##    80% -> 2183 words
##    90% -> 7062 words
##   100% -> 71837 words
## --- NEWS ---
## Words for 50% coverage: 220 - word: health 
## Words for 90% coverage: 8580 - word: stephens 
## 
## Progressive coverage:
##    10% -> 3 words
##    20% -> 10 words
##    30% -> 29 words
##    40% -> 78 words
##    50% -> 220 words
##    60% -> 561 words
##    70% -> 1311 words
##    80% -> 3026 words
##    90% -> 8580 words
##   100% -> 70677 words
## --- TWITTER ---
## Words for 50% coverage: 138 - word: down 
## Words for 90% coverage: 5894 - word: bennett 
## 
## Progressive coverage:
##    10% -> 5 words
##    20% -> 14 words
##    30% -> 34 words
##    40% -> 68 words
##    50% -> 138 words
##    60% -> 290 words
##    70% -> 665 words
##    80% -> 1748 words
##    90% -> 5894 words
##   100% -> 69303 words

7. Q4: Foreign Language Estimation

49K-53K rare words potentially non-English per dataset.

## --- BLOGS ---
## Total unique words: 71837 
## Rare words (freq <= 3): 51255 ( 71.35 % of vocab)
## Estimated foreign/rare: 51255 
## Sample:
## mcbrides, impresses, talentsfor, timeswe, barcis, filings, nudist, jennis, margo, spotlightsprouting, uptrend, punctuality
## --- NEWS ---
## Total unique words: 70677 
## Rare words (freq <= 3): 49135 ( 69.52 % of vocab)
## Estimated foreign/rare: 49135 
## Sample:
## minddestroying, harriman, twittercomrayrinaldi, vickerson, bahr, empties, outspsortcom, hosed, merin, terese, swansons, reynoso
## --- TWITTER ---
## Total unique words: 69303 
## Rare words (freq <= 3): 53730 ( 77.53 % of vocab)
## Estimated foreign/rare: 53730 
## Sample:
## metzger, pleaz, summermusic, tempter, canitbemondayalready, magicians, nunthing, queenstown, mebeen, yella, sonw, aut

8. Q5: Coverage Improvement Strategies

Removing hapax cuts 50% vocab, loses only 2-3% of tokens.

## 1 - Remove hapax (words appearing once):
##   BLOGS: 37064 hapax (51.6% vocab), 2.12% of tokens
##   NEWS: 34292 hapax (48.5% vocab), 2.12% of tokens
##   TWITTER: 42287 hapax (61.0% vocab), 3.04% of tokens
## 
## 2 - Stemming: group 'running/runs/ran' -> 'run' (~30-40% vocab reduction)
## 
## 3 - Reduced dictionary:
##   BLOGS: 7062=90%, 16521=95%, 54331=99% coverage
##   NEWS: 8580=90%, 18614=95%, 54528=99% coverage
##   TWITTER: 5894=90%, 15389=95%, 55384=99% coverage
## 
## 4 - Normalize numbers -> <NUM>, URLs -> <URL>

9. Visualizations

Contexto para os 3 gráficos (barras, cobertura, Zipf).

Top 20 Words - Bar Charts

Dictionary Coverage Curve

Zipf Distribution (Log-Log)

10. Summary Table

~70K unique words per dataset; 6K-8.5K covers 90%.

##              Metric      Blogs       News    Twitter
##        Sample lines   44917.00   50497.00  118272.00
##       Sample tokens 1750622.00 1614933.00 1391912.00
##        Unique words   71837.00   70677.00   69303.00
##          Hapax (1x)   37064.00   34292.00   42287.00
##    Hapax % of vocab      51.59      48.52      61.02
##  Words for 50% cov.     124.00     220.00     138.00
##  Words for 90% cov.    7062.00    8580.00    5894.00
##  Est. foreign words   51255.00   49135.00   53730.00

11. Next Steps: Prediction Algorithm & Shiny App

Based on the findings from this exploratory analysis, I have defined a clear plan for building the prediction algorithm and the interactive application.

Prediction Model

I will build a 4-gram model — meaning it will look at up to three previous words to predict the next one. The model will use Stupid Backoff scoring, an efficient technique for handling word combinations that were not seen in the training data. For rare n-grams, Katz backoff with discounting will smooth the probability estimates. The model will be stored using data.table objects for fast lookups. To keep the model under 100MB of RAM — essential for running on Shiny’s servers — I will apply stemming and remove words that appear only once, which reduces vocabulary by over 50% while losing only 2-3% of token coverage. Profanity and extremely rare words will also be filtered out.

Shiny Application

The final product will be a simple, fast web application where:

  • Users type a phrase into a text box
  • The app suggests the top 3 most likely next words in real time
  • The model loads once when the app starts
  • Response time is targeted at under 500 milliseconds
  • The interface will be clean and minimal, working well on both mobile and desktop devices

Evaluation Strategy

To measure how well the model performs, I will:

  • Set aside 20% of the Twitter data as a test set
  • Measure accuracy — how often the correct word appears in the top 1 and top 3 suggestions
  • Compare different model sizes and speed across various n-gram orders
  • Use perplexity as an additional quality metric
  • Profile memory usage and processing time to ensure smooth deployment on shinyapps.io