Project Overview

The goal of this project is to explore a large collection of English text and understand its main characteristics.

The data comes from three different sources: blogs, news, and Twitter.

The final goal is to build a simple application that can predict the next word based on a phrase entered by the user.

Data Preparation

library(stringr)
library(ggplot2)

blogs <- readLines("data/en_US.blogs.txt", n = 10000)
news <- readLines("data/en_US.news.txt", n = 10000)
twitter <- readLines("data/en_US.twitter.txt", n = 10000)

data_summary <- data.frame(
  Source = c("Blogs", "News", "Twitter"),
  Lines = c(
    length(blogs),
    length(news),
    length(twitter)
  ),
  Words = c(
    sum(str_count(blogs, "\\S+")),
    sum(str_count(news, "\\S+")),
    sum(str_count(twitter, "\\S+"))
  )
)

all_text <- c(blogs, news, twitter)

all_text <- tolower(all_text)

all_text <- gsub("[^a-zA-Z0-9' ]", "", all_text)

words <- unlist(str_split(all_text, "\\s+"))

words <- words[words != ""]

word_freq <- sort(table(words), decreasing = TRUE)

top_words <- head(word_freq, 15)

top_words_df <- data.frame(
  Word = names(top_words),
  Frequency = as.numeric(top_words)
)

Data Summary

The dataset contains three different sources of English text. The table below shows the number of lines and words in each source.

data_summary
##    Source Lines  Words
## 1   Blogs 10000 410620
## 2    News 10000 343929
## 3 Twitter 10000 127674

The three sources contain different types of language. Blogs and news generally contain longer and more structured text, while Twitter contains shorter and more informal messages.

Most Frequent Words

To understand the text better, we counted how often each word appears in the dataset.

head(word_freq, 15)
## words
##   the    to   and     a    of    in     i  that    is   for    it    on  with 
## 43747 23834 22639 21231 18822 14623 13103  9493  9295  9077  7932  6790  6492 
##   you   was 
##  6470  5917

Most Frequent Words Plot

The plot below shows the 15 most frequently used words in the sample.

ggplot(
  top_words_df,
  aes(x = reorder(Word, Frequency), y = Frequency)
) +
  geom_col() +
  coord_flip() +
  labs(
    title = "Top 15 Most Frequent Words",
    x = "Word",
    y = "Frequency"
)

Initial Findings

The exploratory analysis shows that the three sources contain a large amount of English text with different writing styles.

Some common words appear very frequently across the dataset. The repeated occurrence of words and word combinations suggests that the data can be used to identify patterns for next word prediction.

The different sources also provide a mix of formal and informal language, which can help the final prediction model handle different types of user input.

Prediction Algorithm Plan

The next step is to build a prediction algorithm based on word patterns in the training data.

The algorithm will look at the words entered by the user and search for similar word combinations that appeared frequently in the training data.

For example, if a user enters:

“I want to”

the algorithm will look for common words that appeared after this phrase and predict the most likely next word.

The model will use word combinations such as two-word and three-word sequences to improve the prediction.

Shiny Application Plan

The final prediction system will be implemented as a Shiny web application.

The application will provide a simple text box where users can enter a phrase. After the user submits the phrase, the application will return the predicted next word.

The goal is to keep the interface simple and easy to understand for non-technical users.

Conclusion

The exploratory analysis confirms that the available text data contains useful information about word frequency and word patterns.

The next stage is to use these patterns to create the prediction algorithm and integrate it into a Shiny application.

The final application will provide a simple way for users to enter text and receive a predicted next word.