There are two data sources which can be used:
- HC Corpora data set
- Obtained from scraping websites with a web crawler
- Divided into 3 categories by type of website: blogs, news, and Twitter
- Accuracy on a test set is around 12%
- Microsoft Web Language Model
- Utilizes a large corpus obtained from websites
- Contains 1- to 5-gram tokens