The algorithm is based on data from a corpus called HC Corpora which has texts from twitter, blogs and news (click here to see data).
All numbers and punctuation marks were removed. Then the databases were tokenized and all text was changed to lowercase.
With the latter bi-grams, tri-grams and quad-grams were created and joined to have a complete set of words from the three databases.