1. Project Overview

Next Word Prediction

A Shiny application that predicts the next word from an English phrase.

Goal

Build a simple and interactive next-word prediction system using an n-gram language model.

Technologies: R • Shiny • Stringr

2. Data & Preprocessing

Training Data

The model uses samples from three English text sources:

  • Blogs
  • News
  • Twitter

Preprocessing

  1. Convert text to lowercase
  2. Remove unnecessary punctuation
  3. Tokenize text into words
  4. Generate bigrams and trigrams
  5. Count word frequencies

3. Prediction Algorithm

N-Gram Backoff Model

The application uses the last two words of the input phrase to search for a likely next word.

Prediction process

Step 1: Try a trigram based on the last two words.

Step 2: If no match exists, use a bigram based on the last word.

Step 3: If no match exists, return the most frequent word in the corpus.

The most frequently observed candidate is selected.

4. Shiny Application

How It Works

The user enters an English phrase and clicks Predict Next Word.

Example

Input:

I want to

Prediction:

be

The application provides the prediction interactively through a simple Shiny interface.

5. Results & Future Improvements

Current Results

The model successfully generates predictions for different English phrases.

Input Prediction
I want to be
In the world
We are going
This is a
People are going

Future Improvements

  • Train on a larger corpus
  • Improve text preprocessing
  • Add model accuracy evaluation
  • Explore higher-order n-gram models
  • Improve the user interface