Author: Sonja Sahebzad
September 25, 2026
A responsive English word-completion product.
Built in R. Published with Shiny. Measured honestly.
Five slides. Use the navigation controls or arrow keys.
3.21 million training lines from the official English blogs, news and Twitter corpus.
Up to four recent words guide a five-gram model. Kneser-Ney smoothing blends shorter patterns when evidence is limited.
Pruned vocabulary and followers reduce memory. Sparse scoring preserves the compact model’s exact top suggestions.
CPU only.
No external prediction API.
Model loaded once per process.
Suggestions refresh as you continue. A fallback keeps unfamiliar contexts usable.
The Results tab explains measured quality. The How to use tab documents the product.
900 reserved examples. 300 each from blogs, news and Twitter.
Chosen on 600 separate development cases, then frozen before testing.
153 / 900
95% interval 14.7% to 19.6%
257 / 900
95% interval 25.7% to 31.6%
95th percentile 0.83 ms
3,000 warm local calls.
Network and typing delay excluded.
900 / 900 cases received a suggestion. This is availability, not 100% accuracy.
Approximate Wilson intervals describe test uncertainty. Individual confidence is not calibrated. Timing: 600 development inputs, five rounds; exact indexed predictions.
A working web product with fast local computation, clickable suggestions, readable documentation and reproducible evidence.
Short context can miss sentence meaning. New names and rare phrases remain difficult. A modest profanity blocklist is not a complete filter.
Compare longer context under the same hosting budget. Calibrate on separate data. Measure suggestion acceptance with users.
Sources: Coursera SwiftKey corpus; Chen & Goodman, 1996; Posit Shiny.
Code, test protocol, case metrics and full report. Author: Sonja Sahebzad.