The underlying predictive model was constructed using the HC Corpora dataset (Blogs, News, and Twitter entries).
- Text Preprocessing Pipeline:
- Lowercased all text and stripped punctuation, numbers, and extra whitespaces.
- Cleaned out profanity using custom lexical filters.
- N-Gram Tokenization & Back-off Strategy:
- Built frequency tables for Unigrams, Bigrams, and Trigrams.
- Implemented the Stupid Back-off Algorithm:
\[ S(w_i \mid w_{i-k+1}^{i-1}) = \begin{cases} \frac{f(w_{i-k+1}^i)}{f(w_{i-k+1}^{i-1})} & \text{if } f(w_{i-k+1}^i) > 0 \\ \alpha \cdot S(w_i \mid w_{i-k+2}^{i-1}) & \text{otherwise} \end{cases} \]
- Uses an absolute back-off penalty constant (\(\alpha = 0.4\)).