1. Introduction

This report explores three English text sources—blogs, news, and Twitter posts—as an initial step toward a next-word prediction application. It summarizes the datasets, compares text lengths, and identifies frequent words. The purpose is to demonstrate familiarity with the data and outline a practical plan for a prediction algorithm and a Shiny application.

2. Load the data

Place en_US.blogs.txt, en_US.news.txt, and en_US.twitter.txt in the same folder as this report. To keep this first exploratory run manageable in a limited online workspace, the report reads a sample of up to 50,000 lines from each file. The sample is used for the tables and plots below; the source file sizes are reported separately.

files <- c(
  Blogs = "en_US.blogs.txt",
  News = "en_US.news.txt",
  Twitter = "en_US.twitter.txt"
)

missing_files <- files[!file.exists(files)]
if (length(missing_files) > 0) {
  stop(
    "Missing file(s): ", paste(missing_files, collapse = ", "),
    ". Upload all three text files to the same folder as this report."
  )
}

sample_n <- 5000L
text_data <- lapply(files, function(f) {
  readLines(f, n = sample_n, warn = FALSE, encoding = "UTF-8")
})

file_sizes <- file.info(files)$size

3. Basic summary of the three files

The table reports each source’s original file size and the number of lines included in the exploratory sample. Word counts are calculated on the sampled lines, so they are intentionally described as sample word counts rather than full-file totals.

count_words <- function(x) {
  if (length(x) == 0) return(0L)
  sum(lengths(strsplit(trimws(x), "\\\\s+")), na.rm = TRUE)
}

summary_table <- data.frame(
  Dataset = names(text_data),
  Sample_Lines = vapply(text_data, length, integer(1)),
  Sample_Words = vapply(text_data, count_words, numeric(1)),
  Original_File_Size_MB = round(as.numeric(file_sizes) / (1024^2), 1),
  row.names = NULL
)

knitr::kable(summary_table, caption = "Summary of the text sources (word counts are based on up to 50,000 sampled lines)")
Summary of the text sources (word counts are based on up to 50,000 sampled lines)
Dataset Sample_Lines Sample_Words Original_File_Size_MB
Blogs 5000 5000 200.4
News 5000 5000 196.3
Twitter 5000 5000 159.4

The sample allows the main structure of each source to be reviewed without processing every line in the full files. Because sample sizes are capped, these word counts should not be interpreted as full-dataset totals.

4. Text length analysis

This section compares the number of words per sampled line. Histograms help show whether the sample contains mostly short posts or longer passages.

line_word_counts <- lapply(text_data, function(x) {
  if (length(x) == 0) return(integer(0))
  as.integer(lengths(strsplit(trimws(x), "\\\\s+")))
})

length_summary <- data.frame(
  Dataset = names(line_word_counts),
  Mean_Words_Per_Line = round(vapply(line_word_counts, mean, numeric(1), na.rm = TRUE), 2),
  Median_Words_Per_Line = vapply(line_word_counts, median, numeric(1), na.rm = TRUE),
  Maximum_Words_Per_Line = vapply(line_word_counts, max, numeric(1), na.rm = TRUE),
  row.names = NULL
)

knitr::kable(length_summary, caption = "Text length statistics for sampled lines")
Text length statistics for sampled lines
Dataset Mean_Words_Per_Line Median_Words_Per_Line Maximum_Words_Per_Line
Blogs 1 1 1
News 1 1 1
Twitter 1 1 1
old_par <- par(mfrow = c(3, 1), mar = c(4, 4, 3, 1))
for (nm in names(line_word_counts)) {
  vals <- line_word_counts[[nm]]
  if (length(vals) > 0) {
    upper <- as.numeric(quantile(vals, 0.99, na.rm = TRUE))
    hist(
      vals[vals <= upper],
      breaks = 25,
      main = paste("Words per line:", nm),
      xlab = "Words per line (up to 99th percentile)",
      col = "lightblue",
      border = "white"
    )
  }
}

par(old_par)

For readability, the histogram display omits the longest one percent of sampled lines. Those lines remain included in the summary statistics.

5. Frequent words

The following first-pass analysis counts alphabetic words in lowercase text. A short list of very common English words is excluded so that other frequent terms are easier to see. This is a descriptive summary, not a prediction model.

stop_words <- c(
  "the","and","for","that","you","with","was","this","have","are","but",
  "not","they","his","her","she","him","from","all","one","had","has",
  "were","what","when","your","will","can","would","there","their","about",
  "our","out","who","its","just","like","get","got","been","into","than",
  "then","them","because","some","more","how","if","or","as","at","to","of",
  "in","on","is","it","a","an","be","by","i","we","he","me","my","so","do"
)

get_top_words <- function(x, n = 10) {
  text <- tolower(paste(x, collapse = " "))
  tokens <- unlist(regmatches(text, gregexpr("[a-z]+", text, perl = TRUE)))
  tokens <- tokens[nzchar(tokens) & !tokens %in% stop_words]
  if (length(tokens) == 0) {
    return(data.frame(Word = character(), Frequency = integer()))
  }
  counts <- sort(table(tokens), decreasing = TRUE)
  top <- head(counts, n)
  data.frame(Word = names(top), Frequency = as.integer(top), row.names = NULL)
}

top_words <- lapply(text_data, get_top_words)
for (nm in names(top_words)) {
  cat("\n\n### ", nm, "\n\n", sep = "")
  print(knitr::kable(top_words[[nm]], caption = paste("Ten frequent words in the", nm, "sample")))
  cat("\n\n")
}

Blogs

Ten frequent words in the Blogs sample
Word Frequency
s 1920
t 1141
up 622
time 519
m 461
no 386
which 373
know 366
now 359
new 341

News

Ten frequent words in the News sample
Word Frequency
s 2327
said 1241
t 652
year 447
up 392
two 340
new 333
after 312
which 294
m 287

Twitter

Ten frequent words in the Twitter sample
Word Frequency
s 711
t 456
m 289
up 271
love 241
good 221
day 213
thanks 200
rt 189
don 170
old_par <- par(mfrow = c(3, 1), mar = c(5, 8, 3, 1))
for (nm in names(top_words)) {
  df <- top_words[[nm]]
  if (nrow(df) > 0) {
    df <- df[order(df$Frequency), ]
    barplot(
      df$Frequency,
      names.arg = df$Word,
      horiz = TRUE,
      las = 1,
      main = paste("Frequent words:", nm),
      xlab = "Frequency in sample",
      col = "steelblue",
      cex.names = 0.75
    )
  }
}

par(old_par)

6. Initial findings to review

Review the generated tables and plots for these points:

The report deliberately avoids inventing numerical findings. The observations should be based on the tables and plots generated from the files in this workspace.

7. Plan for the prediction algorithm

The planned next steps are to clean and tokenize the text, build tables of one-word, two-word, and three-word sequences (n-grams), and use the most recent words in a phrase to rank possible next words. The model will be evaluated on held-out text not used during training. Accuracy and response speed will both be considered.

8. Plan for the Shiny application

The Shiny application will provide a simple text box for entering a phrase and display a short list of likely next words. It will be tested for useful suggestions, quick response time, and sensible behavior when no matching phrase is available. This report documents exploratory analysis and plans; it does not claim that the prediction model or app is already complete.

9. Conclusion

This report checks the three text sources and explores sample sizes, word counts, line lengths, and frequent words using tables and plots. The sample-based approach makes the initial analysis more manageable while preserving a clear path toward a next-word prediction algorithm and Shiny application.


Before publishing: Confirm the generated tables and plots, then publish the knitted HTML report to RPubs. The word counts and line-length statistics in this report are based on up to 50,000 lines per source, not the complete files.