title: "Next-Word Prediction for Smarter Typing"author: "POOJA AJITH KUMAR"date: "October 2026"output: slidy_presentation

Slide 1: Title & Problem

Next‑Word Prediction for Smarter Typing

A trigram language model with backoff (MCA Capstone)

  • Problem: Predict the next word a user will type, like smartphone keyboards.
  • Goal: Build a fast, accurate next-word predictor from large English text corpora.
  • Output: A Shiny app that takes a phrase and returns predicted next words.

Think: a minimal smart‑keyboard engine you can extend and deploy.

Slide 2: Data & Algorithm

How the Model Works

  • Trained on large English corpora (blogs, news, Twitter).
  • Preprocessing: lowercasing, tokenization, basic cleaning.

Model: trigram language model with backoff

  • Try trigram: P(next | word₋₂, word₋₁)
  • If no match, back off to bigram, then unigram.
  • Only trigrams occurring ≥ 2 times are kept (balances accuracy and size).

Result: Simple, interpretable, and very fast predictions.

Slide 3: Predictive Performance

How Well Does It Predict?

Evaluated on a held‑out test set.

plot of chunk unnamed-chunk-1

  • Top‑1 accuracy: 10.74%
  • Top‑5 accuracy: 23.28%
  • Average prediction time: 0.69 ms per query

Interpretation: With huge vocabularies and short context, top‑5 accuracy > 20% is practically useful for suggestion lists.

Slide 4: The Data Product (Shiny App)

The App: Live Next‑Word Prediction

  • Built with R Shiny.
  • User types a phrase (e.g., I want to) in a text box.
  • Model returns:
    • Top predicted next word.
    • Top‑5 suggestions with probabilities (if implemented).

Live app: https://pooja67.shinyapps.io/nextword_app_token/

How to use:

  1. Open the app link.
  2. Type a phrase in the input box.
  3. Press Predict → see suggested next words.

Slide 5: Why This Is Awesome

Why This Matters & What’s Next

Use cases

  • Faster typing on mobile and web.
  • Assistive technology for users with motor or language difficulties.

Strengths

  • Simple, interpretable model.
  • Very fast predictions (< 1 ms).

Future improvements

  • Smoother probability estimates (e.g., Kneser–Ney smoothing).
  • Neural language models for better context understanding.
  • Domain‑specific tuning (medical, legal, tech, etc.).

This is a minimal but working smart‑keyboard engine that could be extended into a production feature.