Author: Sonja Sahebzad
Utrecht, the Netherlands | September 27, 2026
Twelve selectable languages.
A compact R app with measured results and a clear user guide.
Open the illustrated user guide
Sonja Projects · Version 1.4
English is unchanged. Additional languages are experimental.
Open the live app
English, German, Finnish and Russian use course-corpus text. Eight added languages use Tatoeba example sentences.
A five-gram model uses up to four recent units. Kneser-Ney smoothing blends longer and shorter patterns when evidence is sparse.
Chinese uses ICU word segmentation. Korean retains space-delimited units. Accents, Cyrillic and Hindi vowel signs are preserved.
Lazy loading and a bounded cache.
No external prediction API.
Methods: Chen & Goodman; ICU. Additional data: Tatoeba contributors, CC BY 2.0 FR.
Chinese does not require spaces. Choose 5, 10 or 20 words. Extra buttons offer short endings from matching training phrases.
Open the illustrated user guide
The User guide / Handleiding button explains both features. Short endings remain experimental; the accuracy charts measure the word model.
First: 153/900
Within three: 257/900
Whiskers show 95% sampling intervals. They do not express certainty for one suggestion.
Archived median: 0.58 ms for three suggestions. Network and display add time.
Archived English evaluation; test examples were not reopened. All 900 received a suggestion. This is availability, not 100% accuracy. Individual confidence is not calibrated.
Held-out Tatoeba sentences. Settings were fixed before evaluation.
Different data, training sizes and word units: these are separate results, not a language ranking.
Better representative data, native-speaker review and a compact model using more sentence context.
Lines show Wilson 95% sampling intervals. Similar sentence templates may remain; everyday-message performance is untested. Full results, earlier prototypes, timing and instructions · Data credits. Sonja Sahebzad · Utrecht, the Netherlands.