Unraveling Literary Success - A Comparative NLP Analysis of Popular and Unpopular Science Fiction Novels

Ubaydul H Sami


Department of Computer Science, Lehman College, CUNY

Introduction

We use Natural Language Processing (NLP) to analyze the reasons behind the varying levels of popularity of two science fiction novels. Using download counts from Project Gutenberg to gauge popularity, we compare The Time Machine by H.G. Wells, a widely acclaimed classic with over 7,000 downloads in the past 30 days, to The Planet Mappers by E. Everett Evans, a lesser-known work with only 92 downloads.

The Time Machine novel follows a Victorian scientist who invents a device to travel through time. In the year 802,701 AD, he discovers humanity has evolved into two races: the gentle Eloi and the predatory Morlocks. After his time machine is stolen and his Eloi friend Weena dies, he then travels further into the future, witnessing Earth’s gradual decline. Thirteen million years into the future, he finds a cold, desolate Earth. Upon his return, his story is met with skepticism. He departs on another journey and never returns, leaving behind only two mysterious flowers as evidence of his travels.

In The Planet Mappers, Jon, and Jak, along with their parents, embark on a mission to map uncharted planets. During their perilous journey, their father sustains a serious injury, adding urgency and tension to their quest. Despite this setback, the family perseveres, facing challenges together as they navigate through alien landscapes and encounter unexpected dangers. The father’s injury serves as a pivotal moment in the story, highlighting the risks involved in their mission and the teamwork and resilience of the characters as they strive to overcome adversity.

Objectives

  1. Compare linguistic and thematic content of both books.
  2. Identify factors influencing reader reception of sci-fi literature.
  3. Inform future sci-fi writer by providing insights into factors influencing literary recognition.

Methods

To begin, the full text for both book were retrieved from Project Gutenberg. We then focused on the core narrative by removing extraneous content like introductions and copyright information through text cleaning. Next, each sentence was broken down into individual words, a process known as tokenization. Finally, the stemming algorithm was applied to these tokens to reduce words to their root form. These initial preprocessing steps prepare the textual data for further NLP analysis.

Results

Word Distribution

The Planet Mappers book demonstrates consistent word repetition across chapters, focusing on family-related terms, whereas The Time Machine exhibits greater lexical diversity, including varying scientific terms per chapter.

Additionally, In The Planet Mappers the top words exhibit significantly more repetition than in The Time Machine book. In The Time Machine, the top word ‘time’ appears 200 times, in contrast, The Planet Mappers features the words ‘jon’ and ‘jack’ 555 and 431 times, respectively.

The diversity in “The Time Machine” enhances reader engagement, whereas the repetitive nature of “The Planet Mappers” might detract from its reception.

Sentiment Analysis


The Time Machine book is more negative than The Planet Mappers. The positive portions of ‘The Time Machine’ sentiment are mostly less than 50%, whereas ‘The Planet Mappers’ has a higher proportion of positive sentiment. In ‘The Time Machine’ book, all chapters have a stronger negative than positive sentiment, while ‘The Planet Mappers’ maintains a balance of both. This suggests that the negative tension or thrill in ‘The Time Machine’ may contribute to its greater popularity compared to ‘The Planet Mappers’.

Topic Models

It is an unsupervised machine learning technique, aims to extract topics from large text datasets. Using graphs of coherence and exclusivity, we determined the optimal number of topics for both The Time Machine and The Planet Mappers. We found six distinct themes in The Time Machine book, these included the mechanics of time travel, the contrasting lifestyles of the Eloi and Morlocks, the desolate future Earth, and the existance of nature without humanity. Interestingly, “The Planet Mappers” despite boasting a wider vocabulary, revealed only three themes: spaceship elements, leadership and medical concerns, and the dangers of space exploration.

This discrepancy suggests that “The Time Machine” offers a broader examination of scientific themes, potentially contributing to its popularity compared to “The Planet Mappers.”

Conclusion

Our comparative NLP analysis of The Time Machine and The Planet Mappers reveals that The Time Machine owes its popularity to greater lexical diversity, broader narrative complexity, and stronger emotional engagement. These factors enhance reader interest and contribute to its success. Despite having more words, The Planet Mappers covers fewer scientific concepts and lacks the narrative depth found in The Time Machine. This suggests that a rich variety of themes and emotional tension are critical to the success of science fiction novels, as demonstrated by our analysis.

Next Steps

I will conduct a sentence-level sentiment analysis to uncover the underlying emotions beyond just positive and negative sentiments. This deeper analysis will help us understand the nuanced emotional landscape of each book and its impact on reader engagement.

References

Evans, E. Everett. The Planet Mappers. Project Gutenberg, 2015, www.gutenberg.org/ebooks/50682 (Originally published 1955 by Dodd Mead)

Wells, H.G. The Time Machine. Project Gutenberg, 2004, www.gutenberg.org/ebooks/35 (Originally published 1895 by William Heinemann)

Selected Packages used

Roberts, Margaret E., Brandon M. Stewart, and Dustin Tingley. 2019. “stm: An R Package for Structural Topic Models.” Journal of Statistical Software 91 (2): 1–40. https://doi.org/10.18637/jss.v091.i02.
Roberts, Margaret, Brandon Stewart, and Dustin Tingley. 2023. Stm: Estimation of the Structural Topic Model. http://www.structuraltopicmodel.com/.
Robinson, David, and Julia Silge. 2023. Tidytext: Text Mining Using Dplyr, Ggplot2, and Other Tidy Tools. https://github.com/juliasilge/tidytext.
Wickham, Hadley, Winston Chang, Lionel Henry, Thomas Lin Pedersen, Kohske Takahashi, Claus Wilke, Kara Woo, Hiroaki Yutani, and Dewey Dunnington. 2023. Ggplot2: Create Elegant Data Visualisations Using the Grammar of Graphics. https://ggplot2.tidyverse.org.