2024-06-11

Text Analysis in R

We can utilize a couple of packages in R to analyze text for common words as well as their frequency in various documents. One of these packages is called quanteda and we are going to use it to consume the lyrics from 100 songs. We call the set of lyrics a corpus.

Then we will remove the most common words (stop-words) and then analyze the remaining words for their frequency or occurrence across the entire set of songs. This is called the term frequency iverse document frequency.

Term Frequency and inverse document frequency

First we calculate the term frequency (tf) and this is relatively simple. It’s just the number of times a term appears over the total terms in the document.

\[\begin{equation} tf = \frac{Occurrence\ of\ term\ in\ a\ document}{Total\ terms\ in\ document} \end{equation}\]

Next we calculate the inverse document frequency (idf)

\[\begin{equation} idf = log\left(\frac{Total\ documents\ in\ corpus}{Number\ of\ documents\ with\ the\ term}\right) \end{equation}\]

Calculate tf-idf

Then in order to calculate tf-idf we simply multiply both values together.

\[\begin{equation} \text{tf-idf} = tf \times idf \end{equation}\]

\[\begin{equation} \Rightarrow \left( \frac{Term\ Occurrence}{Total\ terms} \right) \times \ log\left(\frac{Total\ documents}{Documents\ with\ term}\right) \end{equation}\]

Creating a Graph Network with Plotly and igraph

# Start with an empty graph
g = make_empty_graph()

# Add vertices to the graph
g = g %>%
  add_vertices(2, color="red") %>%
  add_edges(edges=c(1,2)) %>%
  add_vertices(1, color="green") %>%
  add_edges(edges=c(2,3,1,3)) %>%
  add_vertices(3, color="blue") %>%
  add_edges(edges=c(1,4,1,5))
  
plot(g)

Displaying the graph

Term Frequency

##   feature frequency rank docfreq group
## 1      \\      6797    1      99   all
## 2       ,      3247    2      94   all
## 3       '      1727    3      82   all
## 4       n      1221    4      95   all
## 5     you      1192    5      85   all

Missing ggplot 1

Missing ggplot 2