This report explores three English text sources—blogs, news, and Twitter posts—as an initial step toward a next-word prediction application. It summarizes the datasets, compares text lengths, and identifies frequent words. The purpose is to demonstrate familiarity with the data and outline a practical plan for a prediction algorithm and a Shiny application.
Place en_US.blogs.txt, en_US.news.txt, and
en_US.twitter.txt in the same folder as this report. To
keep this first exploratory run manageable in a limited online
workspace, the report reads a sample of up to 50,000 lines from each
file. The sample is used for the tables and plots below; the source file
sizes are reported separately.
files <- c(
Blogs = "en_US.blogs.txt",
News = "en_US.news.txt",
Twitter = "en_US.twitter.txt"
)
missing_files <- files[!file.exists(files)]
if (length(missing_files) > 0) {
stop(
"Missing file(s): ", paste(missing_files, collapse = ", "),
". Upload all three text files to the same folder as this report."
)
}
sample_n <- 5000L
text_data <- lapply(files, function(f) {
readLines(f, n = sample_n, warn = FALSE, encoding = "UTF-8")
})
file_sizes <- file.info(files)$size
The table reports each source’s original file size and the number of lines included in the exploratory sample. Word counts are calculated on the sampled lines, so they are intentionally described as sample word counts rather than full-file totals.
count_words <- function(x) {
if (length(x) == 0) return(0L)
sum(lengths(strsplit(trimws(x), "\\\\s+")), na.rm = TRUE)
}
summary_table <- data.frame(
Dataset = names(text_data),
Sample_Lines = vapply(text_data, length, integer(1)),
Sample_Words = vapply(text_data, count_words, numeric(1)),
Original_File_Size_MB = round(as.numeric(file_sizes) / (1024^2), 1),
row.names = NULL
)
knitr::kable(summary_table, caption = "Summary of the text sources (word counts are based on up to 50,000 sampled lines)")
| Dataset | Sample_Lines | Sample_Words | Original_File_Size_MB |
|---|---|---|---|
| Blogs | 5000 | 5000 | 200.4 |
| News | 5000 | 5000 | 196.3 |
| 5000 | 5000 | 159.4 |
The sample allows the main structure of each source to be reviewed without processing every line in the full files. Because sample sizes are capped, these word counts should not be interpreted as full-dataset totals.
This section compares the number of words per sampled line. Histograms help show whether the sample contains mostly short posts or longer passages.
line_word_counts <- lapply(text_data, function(x) {
if (length(x) == 0) return(integer(0))
as.integer(lengths(strsplit(trimws(x), "\\\\s+")))
})
length_summary <- data.frame(
Dataset = names(line_word_counts),
Mean_Words_Per_Line = round(vapply(line_word_counts, mean, numeric(1), na.rm = TRUE), 2),
Median_Words_Per_Line = vapply(line_word_counts, median, numeric(1), na.rm = TRUE),
Maximum_Words_Per_Line = vapply(line_word_counts, max, numeric(1), na.rm = TRUE),
row.names = NULL
)
knitr::kable(length_summary, caption = "Text length statistics for sampled lines")
| Dataset | Mean_Words_Per_Line | Median_Words_Per_Line | Maximum_Words_Per_Line |
|---|---|---|---|
| Blogs | 1 | 1 | 1 |
| News | 1 | 1 | 1 |
| 1 | 1 | 1 |
old_par <- par(mfrow = c(3, 1), mar = c(4, 4, 3, 1))
for (nm in names(line_word_counts)) {
vals <- line_word_counts[[nm]]
if (length(vals) > 0) {
upper <- as.numeric(quantile(vals, 0.99, na.rm = TRUE))
hist(
vals[vals <= upper],
breaks = 25,
main = paste("Words per line:", nm),
xlab = "Words per line (up to 99th percentile)",
col = "lightblue",
border = "white"
)
}
}
par(old_par)
For readability, the histogram display omits the longest one percent of sampled lines. Those lines remain included in the summary statistics.
The following first-pass analysis counts alphabetic words in lowercase text. A short list of very common English words is excluded so that other frequent terms are easier to see. This is a descriptive summary, not a prediction model.
stop_words <- c(
"the","and","for","that","you","with","was","this","have","are","but",
"not","they","his","her","she","him","from","all","one","had","has",
"were","what","when","your","will","can","would","there","their","about",
"our","out","who","its","just","like","get","got","been","into","than",
"then","them","because","some","more","how","if","or","as","at","to","of",
"in","on","is","it","a","an","be","by","i","we","he","me","my","so","do"
)
get_top_words <- function(x, n = 10) {
text <- tolower(paste(x, collapse = " "))
tokens <- unlist(regmatches(text, gregexpr("[a-z]+", text, perl = TRUE)))
tokens <- tokens[nzchar(tokens) & !tokens %in% stop_words]
if (length(tokens) == 0) {
return(data.frame(Word = character(), Frequency = integer()))
}
counts <- sort(table(tokens), decreasing = TRUE)
top <- head(counts, n)
data.frame(Word = names(top), Frequency = as.integer(top), row.names = NULL)
}
top_words <- lapply(text_data, get_top_words)
for (nm in names(top_words)) {
cat("\n\n### ", nm, "\n\n", sep = "")
print(knitr::kable(top_words[[nm]], caption = paste("Ten frequent words in the", nm, "sample")))
cat("\n\n")
}
| Word | Frequency |
|---|---|
| s | 1920 |
| t | 1141 |
| up | 622 |
| time | 519 |
| m | 461 |
| no | 386 |
| which | 373 |
| know | 366 |
| now | 359 |
| new | 341 |
| Word | Frequency |
|---|---|
| s | 2327 |
| said | 1241 |
| t | 652 |
| year | 447 |
| up | 392 |
| two | 340 |
| new | 333 |
| after | 312 |
| which | 294 |
| m | 287 |
| Word | Frequency |
|---|---|
| s | 711 |
| t | 456 |
| m | 289 |
| up | 271 |
| love | 241 |
| good | 221 |
| day | 213 |
| thanks | 200 |
| rt | 189 |
| don | 170 |
old_par <- par(mfrow = c(3, 1), mar = c(5, 8, 3, 1))
for (nm in names(top_words)) {
df <- top_words[[nm]]
if (nrow(df) > 0) {
df <- df[order(df$Frequency), ]
barplot(
df$Frequency,
names.arg = df$Word,
horiz = TRUE,
las = 1,
main = paste("Frequent words:", nm),
xlab = "Frequency in sample",
col = "steelblue",
cex.names = 0.75
)
}
}
par(old_par)
Review the generated tables and plots for these points:
The report deliberately avoids inventing numerical findings. The observations should be based on the tables and plots generated from the files in this workspace.
The planned next steps are to clean and tokenize the text, build tables of one-word, two-word, and three-word sequences (n-grams), and use the most recent words in a phrase to rank possible next words. The model will be evaluated on held-out text not used during training. Accuracy and response speed will both be considered.
The Shiny application will provide a simple text box for entering a phrase and display a short list of likely next words. It will be tested for useful suggestions, quick response time, and sensible behavior when no matching phrase is available. This report documents exploratory analysis and plans; it does not claim that the prediction model or app is already complete.
This report checks the three text sources and explores sample sizes, word counts, line lengths, and frequent words using tables and plots. The sample-based approach makes the initial analysis more manageable while preserving a clear path toward a next-word prediction algorithm and Shiny application.
Before publishing: Confirm the generated tables and plots, then publish the knitted HTML report to RPubs. The word counts and line-length statistics in this report are based on up to 50,000 lines per source, not the complete files.