Word Prediction: Probabilistic Guessing!

Gary De Young
Apr 15, 2016

Language Structure

The English language is filled with structure. This structure is some how mastered by the minds of developing children and their ability grows over time to produce truly complex sentence and paragraphs!

Artificially reproducing this structure in a coherent meaningful way is very difficult; however, some of the structure in language is easliy amenable to computers. It is the structure of word and phrase frequency. Consider the phrase “I'm going to the …” how would you complete it? Perhaps with gym, store, or Colosseum?!

The first two are far more common in most English context

Word prediction

The essence of word prediction is to determine the next likely word.

Implementation

We form from a large body of example text, the 4, 3, an 2-grams (N words in a row within a sentence) and use these to predict a next likely word. Probable words are collected and ordered as follows:

The last word of any 4-gram where the previous 3 words match, is added to a word list that is ranked by estimated probability of occurance. This is repeated with 3-grams, with a two word match, and if needed with a 2-gram, with a single word match, until the requested number of likely words is filled.

Implementation (cont)

We used 384,000 examples of English from blogs, news and twitter (split equally). Spelling checks and other cleaning was applied after N-gram collection so that no false N-grams were added. The N-grams counts were reduced to increase performance speed. A N-gram was included in the prediction set if it met a frequency requirement or a last word length requirement given below.

N-gram Initial Count Freq Restriction Length Inclusion Final Count
Words 85858 None Any 85858
2-gram 1914368 >15 times Any last word 63102
3-gram 4551171 >1 Times Last word > 5 char 675549
4-gram 5475688 >1 Times Last word > 5 char 326943

Prediction Success Rate

The algorithm was tested on random sample of 1000 4-grams formed from a non-training set of English text. The results are below as the frequency of the 4th word appearing in the top predicted words to follow the first 3 words. The success is noticeably improved if the first character of the next word is included.

Character of next wordFirst WordIn top 5In top 10In top 20
0 15.5% 30% 37.7% 43.7%
1 26.4% 51.7% 59.7% 70.6%
2 35.1% 62.7% 70.9% 77.7%

Try it!

Try it at https://gdeyoung.shinyapps.io/Word_Prediction/
Also at the link is an app generated poem.