Sonja Next Word

Author: Sonja Sahebzad
September 25, 2026

SONJA PROJECTS
A small assist for
your next sentence.

A responsive English word-completion product.
Built in R. Published with Shiny. Measured honestly.

Open the live app

Five slides. Use the navigation controls or arrow keys.

Common patterns, compact delivery

Learn from real writing

3.21 million training lines from the official English blogs, news and Twitter corpus.

Combine short and long patterns

Up to four recent words guide a five-gram model. Kneser-Ney smoothing blends shorter patterns when evidence is limited.

Keep only useful continuations

Pruned vocabulary and followers reduce memory. Sparse scoring preserves the compact model’s exact top suggestions.

19.2 MiBcompressed model50,000possible next words

CPU only.
No external prediction API.
Model loaded once per process.

Type. Predict. Keep writing.

  1. Enter an English phrase.
  2. Add a space for automatic suggestions, or select Predict next word.
  3. Select a suggestion to append it, or keep typing.

Suggestions refresh as you continue. A fallback keeps unfamiliar contexts usable.

The Results tab explains measured quality. The How to use tab documents the product.

Try the live app

Live Sonja Next Word app showing a phrase, a primary suggestion and two alternatives.

Performance of the deployed model

900 reserved examples. 300 each from blogs, news and Twitter.
Chosen on 600 separate development cases, then frozen before testing.

17.0%

First suggestion correct

153 / 900
95% interval 14.7% to 19.6%

28.6%

Correct within three

257 / 900
95% interval 25.7% to 31.6%

0.58 ms

Median computation

95th percentile 0.83 ms
3,000 warm local calls.
Network and typing delay excluded.

900 / 900 cases received a suggestion. This is availability, not 100% accuracy.

Approximate Wilson intervals describe test uncertainty. Individual confidence is not calibrated. Timing: 600 development inputs, five rounds; exact indexed predictions.

A deployable foundation, with a clear next step

Value delivered

A working web product with fast local computation, clickable suggestions, readable documentation and reproducible evidence.

Limits made visible

Short context can miss sentence meaning. New names and rare phrases remain difficult. A modest profanity blocklist is not a complete filter.

Next investment

Compare longer context under the same hosting budget. Calibrate on separate data. Measure suggestion acceptance with users.

Explore Sonja Next Word

Sources: Coursera SwiftKey corpus; Chen & Goodman, 1996; Posit Shiny.
Code, test protocol, case metrics and full report. Author: Sonja Sahebzad.