Generative AI is rapidly transforming our lives, and advanced video models are at the forefront. In an era where the proliferation of fake videos is a growing concern, understanding public perceptions of these groundbreaking technologies is critical for their responsible development and successful integration into society. I’ve deeply analyzed public reactions to these new models to uncover key insights for AI’s future development. Therefore, the following research questions were selected: ① “What emotions, reactions, and perceptions will people show regarding the realism of AI-generated videos?”(visual and auditory realism of AI-generated videos, rather than the content itself); ② “What are the key public discussions about AI video models and their outputs, and what factors drive them?”; ③ “What practical insights and implications for the direction of AI technology development can be derived from these analysis results?”
YouTube comment data for OpenAI’s Sora (released in February 2024) and Google’s Veo3 (updated and released in May 2025) was directly collected via web crawling and used as the dataset. To answer the research questions, NRC sentiment analysis, word cloud frequency analysis (for both single words and bigrams), and semantic network analysis (including Co-occurrence, Bigram, and Phi-coefficient methods) were utilized.
NRC sentiment analysis results showed that positive emotions were generally high in both models, with Sora appearing more positive than Veo3. Examining daily sentiment changes over two weeks revealed unique characteristics: Veo3 demonstrated relatively more consistent sentiment shifts. These differences likely stem from factors such as the release timing of each model’s videos, video realism(human-likeness level, etc.), public perception of AI-generated videos, and limitations in technological implementation. This was clearly confirmed through word cloud frequency analysis, highlighting distinct keywords and bigrams. In Sora, keywords like “amazing” and “voice mode” were prominent, while Veo3 featured “cooked”, “real people” and “ai slop”. Furthermore, semantic network analysis allowed for a broader and deeper understanding of core topics, issues, interests, and underlying causes by identifying major clusters and keyword associations. Sora’s semantic network primarily focused on creativity and positive potential, with themes such as “creative potential”, “amazing technology”, and “realistic voice” being prominent. This indicates discussions revolved around “artistic potential”, “future expectations”, “audio feature expectations” and “industry impact”. Notably, keywords like ‘real people/real time/real life’ and ‘uncanny valley’ simultaneously showed awe at Sora’s outstanding realism and a subtle discomfort. In contrast, the Veo3 network was dominated by critical perceptions of reality and social concerns. “AI slop”, “fake content”, “uncanny valley”, “job loss” and “dystopia” were core themes. Veo3 comments strongly criticized content quality, using terms like ‘ai slop’, ‘soulless’ and ‘bad acting’. Furthermore, ‘black mirror’ and ‘sentinel island’ expressed deep social and ethical concerns about AI technology leading to ‘dystopian futures’ or ‘uncontrollable risks’. The conflicting appearance of “real people” and “fake content” in Veo3 suggests that high realism can paradoxically highlight a lack of human sensibility and artistry. Significantly, common nodes like “trust”, “creativity”, “ethics” and “future” connected in the semantic networks of both models. This indicates that beyond mere technical perfection, creativity, empathy, and social responsibility are emerging as crucial factors for AI acceptance. Additionally, the ‘jobs/scary’ cluster, present in both models, highlights a shared concern about ‘job displacement’, suggesting that the social implications of technology, beyond just realism, are driving public discourse.
Accordingly, AI technology development should proceed in three directions. First, designing with inherent artistry and creativity is crucial. This means that beyond merely advancing realism, AI-generated content should be designed to satisfy human aesthetics and sensibilities to address issues like the ‘uncanny valley’ and ‘unnaturalness’. Second, development that reflects human emotions is necessary. AI should evolve to understand and express human experiences and emotions, rather than just imitating reality, thereby gaining user ‘trust’ and ‘empathy’ and stabilizing technology adoption. Third, transparent communication regarding AI ethics, social trust, and changes in the job market must be established. Deep reflection and a proactive approach to potential misuse, social repercussions, employment changes, and ethical issues will be essential. These insights pose crucial questions not only for technology developers but also for all users consuming AI content and society as a whole observing this trend. Ultimately, AI video technology must aim for ‘human-centered AI’, prioritizing the artistic and human value of content and considering its societal impact, beyond just functional accuracy.
Public introduction videos for Sora and Veo3 were directly sourced from YouTube. Commentary videos by professional YouTubers were largely excluded to maintain research directness. For Sora, 3,140 comments were collected from 5 videos (IDs: 2fAPgOCJToA, anmukFtu8U, drmNl6-UBmBI, f75eoFyo9ns, y_4Kv_Xy7vs) with a total of 2,138,275 views. For Veo3, 3,765 comments were gathered from 10 videos (IDs: x_x-JAAKSVJ, 94KmlFYIA08, 9rmHjbNhfos, Cqcf3XEWPeo, IOg8QvkKhO1A, QOQnTlwueOpY, clG3yiKfAFtl, 5_CvdCLcfUA, Cxx92BbHhBW, ODYROOW1dCo) totaling 1,389,232 views. Videos with high view and comment counts were selected to ensure a comparable volume of comments for both models.
Web crawling was performed directly using the YouTube API. The collected comment data was saved into Sora2024_comments.csv and Veo3-2025_comments.csv files. The original data included video_id, comment_id, parent_id, author, text, published_at, like_count, with an added date (yyyy-mm-dd) column. Of the total 6,905 comments for both models, most were English text; however, non-English comments (e.g., Nepali, Russian, Arabic) and those consisting solely of emojis required preprocessing.
Sora and Veo3 YouTube comment collection files were loaded as character-type data frames. Using the mutate() function, an ai_model column was assigned “Sora” or “Veo3” to each, and these data frames were combined to create all_comments. The date column in all_comments was then converted to a date type.
A custom clean_text function was applied to the text column, generating preprocessed_comments. This function performed several cleaning steps: lowercasing, removing URLs, emails/hashtags, punctuation/special characters, and non-English characters. During this process, empty comment rows were removed. A comments_code was assigned to each comment to identify its source before tokenization. Finally, tokenized_comments data was built by performing tokenization, removing common stop words, and eliminating purely numeric tokens. The final structure of the data frame can be confirmed as follows.
# Define custom stop words for frequency analysis
custom_stop_words <- bind_rows(tibble(word = c("don","gonna","doesn","isn","ll","ve","didn","dont","wasn","im","de","sora","veo3","veo","openai","google","altman","joe","video","videos","wow","bro","ship","generated","spaghetti","eating","generate","won","lot","shit","yeah","lol","ma","waithard","created","create","technology","ai","people","time","real","human","humans","world","movie","movies","imagine","feel","prompt","prompts","music","future","guy"),
lexicon = c("custom")),
stop_words)
# Define custom2 stop words for bigram analysis
custom2_stop_words <- bind_rows(tibble(word = c("don","gonna","doesn","isn","ll","ve","didn","dont","wasn","im","de","sora","veo3","veo","openai","google","altman","joe","video","videos","wow","bro","ship","generated","spaghetti","eating"),
lexicon = c("custom2")),
stop_words)
# Define custom3 stop words for network analysis
custom3_stop_words <- bind_rows(tibble(word = c("don","gonna","doesn","isn","ll","ve","didn","de","es","en","la","las","los","el","di","una","se","esto","est","wanna","won","bro","wow","lot","shit","yeah","lol","ma","waithard","spaghetti","eating","piano","generated","openai","google","video","videos","created","create","technology","ai","people","human","humans","world","music","movie","movies","prompt","prompts","feel","time","tom","rick","morty","drinks","energy","gpt","intelligence","gamozolas","joe","biden","zaha","hadid","3d","artist","sci","fi","sora","veo","company","companies","generate","film", "fu","images","text","tim","choses","6th","continue","wasn","stories","platforms","jesus","crist","im","dont","forward","step","august","kamp","trek","smith","pretty","4o","till","day","produce","written","favorite","dollars","director","couple","real","absolutely","life","stuff"),
lexicon = c("custom3")),
stop_words)
# 1. Data Loading and Merging
Sora_comments <- read.csv("Sora2024_comments.csv")
Sora_comments <- Sora_comments %>%
mutate(ai_model = "Sora")
Veo3_comments <- read.csv("Veo3-2025_comments.csv")
Veo3_comments <- Veo3_comments %>%
mutate(ai_model = "Veo3")
all_comments <- bind_rows(Sora_comments, Veo3_comments) %>%
mutate(date = as.Date(date, format = "%Y-%m-%d"))
# 2. Text Cleaning Function
clean_text <- function(text) {
text %>%
str_to_lower() %>%
str_replace_all("http\\S+\\s*", " ") %>%
str_replace_all("www\\.\\S+\\s*", " ") %>%
str_replace_all("(@\\S+)", " ") %>%
str_replace_all("(#\\S+)", " ") %>%
str_replace_all("[[:punct:]]", " ") %>%
str_replace_all("[[:cntrl:]]", " ") %>%
str_replace_all("[^a-zA-Z0-9\\s]", " ") %>%
str_squish() }
# 3. Preprocessing Comments
# Apply cleaning function, handle missing values and add comment_code
preprocessed_comments <- all_comments %>%
mutate(cleaned_text = clean_text(text)) %>%
drop_na(cleaned_text) %>%
rowid_to_column("comments_code")
# 4. Tokenization and Filtering
data("stop_words")
tokenized_comments <- preprocessed_comments %>%
unnest_tokens(word, cleaned_text) %>%
anti_join(stop_words, by = "word") %>%
filter(!str_detect(word, "^[0-9]+$")) %>%
select(comments_code, ai_model, word, date)
head(tokenized_comments)
## comments_code ai_model word date
## 1 1 Sora openai 2025-04-24
## 2 1 Sora ko 2025-04-24
## 3 1 Sora bhi 2025-04-24
## 4 1 Sora isiliye 2025-04-24
## 5 1 Sora abhi 2025-04-24
## 6 1 Sora public 2025-04-24
This project’s key variables of interest are:
“ai_model” (a core categorical variable comparing Sora
and Veo3 models, chosen for their innovative first and latest
significance among various AI video generation models),
“comments_code” (a unique identifier for each valid
comment after preprocessing, criterion to recognize individual comments
as independent text units for co-occurrence network analysis),
“word” (individual tokens after preprocessing and stop
word removal, serving as the fundamental unit for text analysis), and
“date” (the comment’s creation date, useful for
analyzing temporal trends). For the “word” variable,
which is central to our text analysis, we utilized the tokenized
tokenized_comments data frame to determine unique word
counts, average words per comment, and identified the most frequently
mentioned words through word frequency analysis.
# Calculate the total number of unique words across all comments
total_unique_words <- tokenized_comments %>%
distinct(word) %>%
nrow()
# Calculate word counts, unique words, and Type-Token Ratio (TTR) by AI model
model_word_counts <- tokenized_comments %>%
group_by(ai_model) %>%
summarise(
Total_Words = n(),
Unique_Words = n_distinct(word)) %>%
ungroup() %>%
mutate(Type_Token_Ratio = Unique_Words / Total_Words)
# Calculate comment length statistics (average, median, standard deviation) by AI model
comment_length_stats <- preprocessed_comments %>%
mutate(Words_Per_Comment = str_count(cleaned_text, "\\S+")) %>%
group_by(ai_model) %>%
summarise(
Avg_Len = mean(Words_Per_Comment, na.rm = TRUE),
Med_Len = median(Words_Per_Comment, na.rm = TRUE),
SD_Len = sd(Words_Per_Comment, na.rm = TRUE)) %>%
ungroup()
# Get all comment word counts for overall stats calculation.
all_comments_word_counts <- preprocessed_comments %>%
mutate(Words_Per_Comment = str_count(cleaned_text, "\\S+")) %>%
pull(Words_Per_Comment)
# Combine model-specific stats and add an overall "Total" row
summary_df <- model_word_counts %>%
left_join(comment_length_stats, by = "ai_model") %>%
add_row(
ai_model = "Total",
Total_Words = sum(.$Total_Words, na.rm = TRUE),
Unique_Words = total_unique_words,
Type_Token_Ratio = total_unique_words / sum(.$Total_Words, na.rm = TRUE),
Avg_Len = mean(all_comments_word_counts, na.rm = TRUE),
Med_Len = median(all_comments_word_counts, na.rm = TRUE),
SD_Len = sd(all_comments_word_counts, na.rm = TRUE)
) %>%
arrange(desc(ai_model == "Total"))
summary_df
## # A tibble: 3 × 7
## ai_model Total_Words Unique_Words Type_Token_Ratio Avg_Len Med_Len SD_Len
## <chr> <int> <int> <dbl> <dbl> <dbl> <dbl>
## 1 Total 41271 9096 0.220 16.3 10 23.4
## 2 Sora 19917 5725 0.287 17.1 10 22.2
## 3 Veo3 21354 6055 0.284 15.5 9 24.4
The analysis revealed that Veo3-related discussions showed a higher
total word count (21,354) compared to Sora
(19,917), indicating more active discourse surrounding
Veo3. Both models demonstrated similarly low levels of lexical
diversity, with Type-Token Ratios (TTR) being close
(Sora: 0.287, Veo3: 0.284). This low TTR means that
approximately 28-29% of the words are unique, with the
rest being repetitions. However, considering the substantial length of
the texts (around 20,000 words each) and the nature of
YouTube comments where core terms are frequently reiterated, this low
TTR is viewed as a natural phenomenon rather than indicating low
quality. Instead, it suggests a concentration of discussions around
specific keywords and themes, which is considered beneficial for the
study’s objective of identifying key insights and discussion points.
Furthermore, comment lengths were relatively short and concise,
averaging 15 to 17 words (Sora Avg_Len: 17.1, Veo3
Avg_Len: 15.5), and Veo3 comments exhibited a slightly larger standard
deviation in word count (SD_Len: 24.4) than Sora (22.2), suggesting
potentially greater variability in comment lengths for Veo3.
# Calculate top 15 most frequent words for all comments
word_counts_total <- tokenized_comments %>%
count(word, sort = TRUE) %>%
head(15) %>%
mutate(rank = row_number()) %>%
rename(Total_Word = word, Total_Count = n)
# Calculate top 15 most frequent words for Sora model comments
word_counts_Sora <- tokenized_comments %>%
filter(ai_model == "Sora") %>%
count(word, sort = TRUE) %>%
head(15) %>%
mutate(rank = row_number()) %>%
rename(Sora_Word = word, Sora_Count = n)
# Calculate top 15 most frequent words for Veo3 model comments
word_counts_Veo3 <- tokenized_comments %>%
filter(ai_model == "Veo3") %>%
count(word, sort = TRUE) %>%
head(15) %>%
mutate(rank = row_number()) %>%
rename(Veo3_Word = word, Veo3_Count = n)
# Combine all three word frequency dataframes into one
combined_word_counts <- full_join(word_counts_total, word_counts_Sora, by = "rank") %>%
full_join(word_counts_Veo3, by = "rank") %>%
select(-rank)
combined_word_counts
## Total_Word Total_Count Sora_Word Sora_Count Veo3_Word Veo3_Count
## 1 ai 1305 ai 575 ai 730
## 2 people 532 video 284 real 282
## 3 video 524 people 258 people 274
## 4 real 440 sora 216 video 240
## 5 videos 252 real 158 veo 170
## 6 time 250 videos 147 time 140
## 7 sora 233 music 122 google 130
## 8 don 230 world 116 don 117
## 9 world 205 don 113 videos 105
## 10 human 174 time 110 human 95
## 11 veo 170 technology 98 world 89
## 12 movie 167 generated 85 lol 87
## 13 generated 163 imagine 83 movie 86
## 14 future 155 future 82 movies 79
## 15 music 154 create 81 generated 78
Analysis of high-frequency words revealed that common keywords across AI video generation technology, such as ‘ai’, ‘video’, ‘real’, and ‘model’, were prominent for both models. Additionally, distinct discussion themes were identified for each model. Sora discussions were characterized by words like ‘sora’ (the model name) alongside ‘music’, ‘imagine’, ‘create’, and ‘future’, suggesting active discourse around the ‘artistic quality’, ‘creativity’, and ‘future potential’ of its outputs. Conversely, Veo3 discussions frequently mentioned ‘google’ (the developer) along with ‘veo’ (the model name), ‘people’, ‘lol’, ‘movie’, ‘long’, and ‘test’, indicating a focus on ‘Google’s association’, ‘user experience evaluations’, and ‘content length’ in practical, specific use cases.
However, the current list of top words includes meaningless terms like ‘don’, ‘ve’, and ‘ll’ (likely derived from “don’t”, “have”, “will” during preprocessing), as well as common and general AI video generation-related words such as ‘ai’ and ‘video’. This limits the ability to conduct a deeper analysis of the core themes in comments for both models. Therefore, it is necessary to remove these words by adding ““custom stop words”“.
Furthermore, words unique to specific videos of each model (e.g., ‘piano’ in comments on a Sora piano video, ‘spaghetti’ in comments on a Veo3 spaghetti video) can also appear among the top words. Since these terms might emerge as distinctive words in TF-IDF analysis, it would be beneficial to include them in the custom stop word list for more in-depth thematic analysis. Nevertheless, for crucial general terms like ‘ai’, ‘video’, and ‘real’, which can form significant meanings when linked in N-grams, performing an N-gram analysis prior to stop word removal would be helpful for a more profound understanding of their contextual meaning.
This analysis visualizes the emotional content of comments related to Sora and Veo3 AI models using the NRC sentiment lexicon, segmenting emotions beyond positive/negative into 8 basic categories like trust and anticipation.
First, tokenized_comments were inner_joined with the NRC lexicon, tagging each word with its corresponding sentiment(s) (allowing for multiple sentiments per word, though this applied to very few words in practice). Then, word percentages for each sentiment were calculated per AI model, preventing misinterpretation from total word count differences. The 10 NRC sentiments were further grouped into ‘Positive’ and ‘Negative’ emotion groups for broader trend identification.
Visualization employed ggplot2. Within each emotion group, the highest percentage sentiments were placed at the top, and Sora and Veo3 bars were positioned side-by-side for easy comparison. The facet_wrap function was used to create separate panels for positive and negative emotion groups, and the x-axis scales were set uniformly so that the relative proportions could be directly read from the length of the bars. Necessary functions and options were implemented with online assistance. This visualization maximized readability to easily discern sentiment differences per model. It effectively aided in quickly identifying the dominant emotional tones and differences across the positive/negative spectrum in comments for both Sora and Veo3.
nrc_sentiment <- get_sentiments("nrc")
sentiment_analysis_nrc <- tokenized_comments %>%
inner_join(nrc_sentiment, by = "word", relationship = "many-to-many")
# Calculate word counts and percentages by AI model and sentiment
sentiment_counts_nrc <- sentiment_analysis_nrc %>%
group_by(ai_model, sentiment) %>%
summarise(word_count = n(), .groups = 'drop') %>%
group_by(ai_model) %>%
mutate(total_words_ai_model = sum(word_count),
word_percentage = (word_count / total_words_ai_model) * 100) %>%
ungroup() %>%
# Classify sentiment into positive/negative emotion groups
mutate(emotion_group = case_when(
sentiment %in% c("positive", "trust", "anticipation", "joy", "surprise") ~ "Positive Emotions",
sentiment %in% c("negative", "fear", "sadness", "anger", "disgust") ~ "Negative Emotions"))
# Set factor levels for sentiment (control graph order)
sentiment_counts_nrc$sentiment <- factor(
sentiment_counts_nrc$sentiment,
levels = c("surprise", "joy", "anticipation", "trust", "positive",
"disgust", "anger", "sadness", "fear", "negative"))
# Set factor levels for emotion groups
sentiment_counts_nrc$emotion_group <- factor(sentiment_counts_nrc$emotion_group,
levels = c("Positive Emotions", "Negative Emotions"))
# Visualize NRC sentiment analysis (percentage of sentiment words by model)
ggplot(sentiment_counts_nrc, aes(x = word_percentage, y = sentiment, fill = ai_model)) +
geom_bar(stat = "identity", position = position_dodge2(reverse = TRUE)) +
facet_wrap(~ emotion_group, scales = "free_y", ncol = 2) +
scale_fill_manual(values = c("Sora" = "#2196F3", "Veo3" = "#FFC107")) +
labs(
title = "NRC Emotion Analysis: Sora vs. Veo3 Comments (by Percentage)",
x = "Word Percentage (%)",
y = "Emotion",
fill = NULL) +
theme_minimal() +
theme(plot.title = element_text(size = 11, hjust = 0.5),
legend.position = "right", strip.text.y = element_text(angle = 0),
axis.title = element_text(size = 9),
axis.text = element_text(size = 9),
strip.text = element_text(size = 9),
legend.text = element_text(size = 9))
Fig.1 NRC Emotion Trends of Sora and Veo3 YouTube Comments
NRC sentiment analysis reveals how AI video realism impacts user experience and emotional responses, providing key insights into technology acceptance. The analysis centered on realism responses, driven by perceptions of visual and auditory implementation (e.g., voice mode, audio) – whether present or enhanced. User comment sentiment for Sora and Veo3 reflects shifts in perceived realism and corresponding user expectations.
Sora’s innovative debut, delivering an “unprecedented new world” experience through its realism, generated overwhelmingly positive user responses. High ‘positive’ and ‘joy’ sentiments indicate users’ awe and satisfaction with its realism and new possibilities. This suggests that at launch, realism’s visual impact and inherent promise drove high technology acceptance, largely overshadowing any technical shortcomings. High ‘anticipation’ also reflected significant positive future expectations.
Conversely, Veo3 showed a lower ‘positive’ and slightly higher overall ‘negative’ sentiment compared to Sora. Despite Veo3’s technical advancements nearing “real AI” realism, users adopted a more critical perspective, considering broader implications and potential issues. Notably, Veo3 exhibited higher ‘trust’ and ‘surprise’ sentiments than Sora, while ‘sadness’ was lower among negative emotions. This suggests that trust in Google as a developer or improved stability influenced user experience and acceptance. The lower ‘sadness’ indicates criticisms are more about technical completeness or societal impact than emotional disappointment, showing how realism-based critical evaluation integrates into user experience as technology matures.
In summary, increased AI video realism initially drives high technology acceptance via ‘surprise’ and ‘anticipation’ (Sora). However, as realism advances towards “real AI” levels (Veo3), users become more critical, actively scrutinizing potential risks and imperfections. This evolving emotional response underscores how technology acceptance extends beyond functional satisfaction, becoming deeply linked with complex experiences like ‘trust’, ‘concern’ and ‘critical evaluation’ as realism deepens.
These graphs show daily emotional trends in comments for Sora and Veo3 AI models. Comment data for two weeks from each model’s launch was analyzed, as comment volume significantly dropped afterwards. From the existing NRC-sentiment-tagged data (sentiment_analysis_nrc), comments within these specific two-week periods were extracted.
Extracted comments were then grouped by AI model, date, and sentiment type. The daily frequency of each sentiment word (word_count) was calculated. Subsequently, the daily percentage (word_percentage) for each sentiment was computed based on the total daily sentiment words per model. This helps avoid misinterpretations from changing comment volumes over time. Sentiments were categorized into positive and negative groups, and their display order was set for clarity.
Using the facet_grid function, a grid graph was created with sentiments as rows and models as columns. Each model’s date axis was set to scale independently, while the y-axis was adjusted to the same scale to facilitate relative comparison. The negative sentiment trend graph was generated separately to avoid excessive complexity and improve readability. To facilitate comparison, the two graphs were arranged horizontally side-by-side.
# Define the start and end dates for ai_model's analysis window (2 weeks from launch)
sora_start_date <- as.Date("2024-02-16")
sora_end_date <- sora_start_date + weeks(2) - days(1)
veo3_start_date <- as.Date("2025-05-21")
veo3_end_date <- veo3_start_date + weeks(2) - days(1)
# Calculate daily sentiment word counts and percentages for the specified periods
combined_daily_sentiment <- sentiment_analysis_nrc %>%
filter(
(ai_model == "Sora" & date >= sora_start_date & date <= sora_end_date) |
(ai_model == "Veo3" & date >= veo3_start_date & date <= veo3_end_date)) %>%
group_by(ai_model, date, sentiment) %>%
summarise(word_count = n(), .groups = 'drop') %>%
group_by(ai_model, date) %>%
mutate(total_words_daily = sum(word_count),
word_percentage = (word_count / total_words_daily) * 100) %>%
ungroup()
# Classify sentiments into broader positive/negative emotion groups
combined_daily_sentiment <- combined_daily_sentiment %>%
mutate(emotion_group = case_when(
sentiment %in% c("positive", "trust", "anticipation", "joy", "surprise") ~ "Positive Emotions",
sentiment %in% c("negative", "fear", "sadness", "anger", "disgust") ~ "Negative Emotions"))
# Set factor levels for individual sentiments to control plotting order
combined_daily_sentiment$sentiment <- factor(
combined_daily_sentiment$sentiment,
levels = c("positive", "trust", "anticipation", "joy", "surprise",
"negative", "fear", "sadness", "anger", "disgust"))
# Set factor levels for emotion groups to control plotting order
combined_daily_sentiment$emotion_group <- factor(
combined_daily_sentiment$emotion_group,
levels = c("Positive Emotions", "Negative Emotions"))
# Plotting Positive Emotions Daily Trends
positive_daily_sentiment <- combined_daily_sentiment %>%
filter(emotion_group == "Positive Emotions")
ggplot(positive_daily_sentiment, aes(x = date, y = word_percentage, fill = ai_model)) +
geom_col(position = "dodge") +
facet_grid(sentiment ~ ai_model, scales = "free_x") +
scale_fill_manual(values = c("Sora" = "#2196F3", "Veo3" = "#FFC107")) +
scale_y_continuous(limits = c(0, 27)) +
labs(
title = "Positive Emotions",
x = "Date",
y = "Word Percentage (%)",
fill = NULL) +
theme_minimal() +
theme(legend.position = "none", strip.text.y = element_text(angle = 0),
axis.title = element_text(size = 16),
axis.text = element_text(size = 13),
strip.text = element_text(size = 16))
# Plotting Negative Emotions Daily Trends
negative_daily_sentiment <- combined_daily_sentiment %>%
filter(emotion_group == "Negative Emotions")
ggplot(negative_daily_sentiment, aes(x = date, y = word_percentage, fill = ai_model)) +
geom_col(position = "dodge") +
facet_grid(sentiment ~ ai_model, scales = "free_x") +
scale_fill_manual(values = c("Sora" = "#2196F3", "Veo3" = "#FFC107")) +
scale_y_continuous(limits = c(0, 27)) +
labs(
title = "Negative Emotions",
x = "Date",
y = "Word Percentage (%)",
fill = NULL) +
theme_minimal() +
theme(legend.position = "none", strip.text.y = element_text(angle = 0),
axis.title = element_text(size = 16),
axis.text = element_text(size = 13),
strip.text = element_text(size = 16),
axis.title.y = element_text(color = "white"),
axis.text.y = element_text(color = "white"))
Fig.2 Daily NRC Emotion Trends of Sora and Veo3 YouTube Comments
Analyzing daily sentiment trends for Sora and Veo3 models reveals that, similar to overall sentiment analysis, positive emotions consistently outweighed negative ones daily. However, daily fluctuations offer more nuanced insights into user engagement, particularly concerning video realism and technology acceptance.
Sora’s sentiment trends show significantly greater daily volatility compared to Veo3, seen in notable drops in ‘trust’ or sharp rises in ‘anger’. This suggests users’ emotional responses to Sora videos, viewed from a broader perspective, exhibited considerable daily shifts. Initial wonder and excitement might have caused emotional swings, reacting to specific video releases or emerging discussions.
In contrast, Veo3 displays a much more consistent sentiment trend, even with comments from 10 different video files. This implies user engagement with Veo3 content was more detailed, critical, and steady regarding video realism and technology acceptance. User reactions were less prone to extreme daily changes, indicating a continuous evaluation of features and implications, likely stemming from a focus on consistent performance. These differences in daily volatility clearly illustrate how users experienced and reacted to the initial offerings of each AI video model.
Key words and themes were aimed to be identified through word frequency analysis in the comments. As can be observed in Fig. 3-1, “AI” was the most frequent word for both models (Sora and Veo3), followed by “video”, “generated”, and “people”. These words were common to both models rather than differentiating them, which limited the prominence of other significant terms. Consequently, it was deemed necessary to treat these words as custom stopwords. A list of custom stopwords was finalized through various simulations, and the resulting word cloud is presented in Fig. 3-2. Other analytical methods, such as TF-IDF and Log-Odds Ratio analysis, were also attempted to identify key keywords. However, these methods often highlighted rare words specific to individual videos for each model, again indicating that optimization via custom stopwords was required. Word cloud analysis was actively utilized because it allowed for easy visual tracking of keyword changes as words were added to the stopword group.
# Process words for Sora Model (common stop words)
word_common_Sora <- tokenized_comments %>%
filter(ai_model == "Sora") %>%
anti_join(stop_words, by = "word") %>%
count(word)
wordcloud(word_common_Sora$word, word_common_Sora$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 1.7))
mtext("Sora Word", side = 3, line = 1, adj = 0.5, cex = 3.2)
# Process words for Veo3 Model (common stop words)
word_common_Veo3 <- tokenized_comments %>%
filter(ai_model == "Veo3") %>%
anti_join(stop_words, by = "word") %>%
count(word)
wordcloud(word_common_Veo3$word, word_common_Veo3$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 1.7))
mtext("Veo3 Word", side = 3, line = 1, adj = 0.5, cex = 3.2)
Fig.3-1 Word Cloud of Sora and Veo3 YouTube Comments: Identifying Key Terms after Common Stop Word Removal
# Process words for Sora Model (custom stop words)
word_custom_Sora <- tokenized_comments %>%
filter(ai_model == "Sora") %>%
anti_join(custom_stop_words, by = "word") %>%
count(word)
wordcloud(word_custom_Sora$word, word_custom_Sora$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 0.3))
mtext("Sora Word", side = 3, line = 1, adj = 0.5, cex = 3.2)
# Process words for Veo3 Model (common stop words)
word_custom_Veo3 <- tokenized_comments %>%
filter(ai_model == "Veo3") %>%
anti_join(custom_stop_words, by = "word") %>%
count(word)
wordcloud(word_custom_Veo3$word, word_custom_Veo3$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 0.3))
mtext("Veo3 Word", side = 3, line = 1, adj = 0.5, cex = 3.2)
Fig.3-2 Word Cloud of Sora and Veo3 YouTube Comments: Identifying Key Terms after Custom Stop Word Removal
As seen in the Sora word cloud on the left side of Fig. 3-2, Key themes in Sora-related comments include: creative / artists / art (creation, artistic potential), music / voice (technical limitations of audio/music, interest in future integration), jobs / scary (job concerns and fear), dream / reality (realistic video, sense of reality), imagine / future (future possibilities, potential for growth), and stock footage / film industry (impact on stock footage market, film industry). While “artistic potential” and “future expectations” were high in Sora comments, concerns about “job displacement” and “reality/fiction confusion” were also identifiable.
In the Veo3 word cloud on the right, prominent keywords include: slop / cooked / fake (strong critical perception, e.g., “AI slop” denoting ‘soulless and meaningless’ output), art / hollywood / acting (interest in impact on art, acting, and Hollywood, often from a critical perspective like ‘low quality’ or ‘fake acting’), money / jobs (impact on content production costs and related employment), quality / content (video quality, content itself), scary / bad (negative emotions), characters / humanity (realism of characters, AI’s portrayal of humanity), and audio / sound (quality of audio features). Veo3 comments, despite mentioning technical completeness (especially audio) compared to Sora, showed strong criticism (“AI slop”) regarding the content’s lack of ‘artistry’ or ‘meaning’. Notably, concerns about “jobs” were common across comments for both models.
To identify the main themes within the comments, I performed a bigram word cloud analysis, filtered by both common and custom stopwords. This approach allowed for easier topic identification compared to simple word clouds, with custom stopwords enabling more detailed theme discovery. Throughout the process of refining custom stopwords, the visual changes in the bigram word cloud were easily observable, enhancing analytical efficiency.
# Process Bigrams for Sora Model
Sora_bigrams <- preprocessed_comments %>%
filter(ai_model == "Sora") %>%
drop_na(cleaned_text) %>%
unnest_tokens(bigram, cleaned_text, token = "ngrams", n = 2) %>%
filter(!is.na(bigram))
# Separate bigrams into individual words to filter out stop words.
Sora_bigrams_separated <- Sora_bigrams %>%
separate(bigram, c("word1", "word2"), sep = " ", remove = FALSE)
# Filter out common stop words and purely numeric words
Sora_bigrams_filtered_common <- Sora_bigrams_separated %>%
filter(!word1 %in% stop_words$word) %>%
filter(!word2 %in% stop_words$word) %>%
filter(!str_detect(word1, "^[0-9]+$")) %>%
filter(!str_detect(word2, "^[0-9]+$"))
# Re-unite the filtered words back into bigrams
Sora_bigrams_united_common <- Sora_bigrams_filtered_common %>%
unite(bigram, word1, word2, sep = " ")
# Count the frequency of each bigram and keep only those appearing 3 or more times
Sora_bigram_counts_common <- Sora_bigrams_united_common %>%
count(bigram, sort = TRUE) %>%
filter(n >= 3)
# Generate and display a word cloud for Sora's bigrams.
wordcloud(Sora_bigram_counts_common$bigram, Sora_bigram_counts_common$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 1))
mtext("Sora Bigram", side = 3, line = 0, adj = 0.5, cex = 3.2)
# Process Bigrams for Veo3 Model
Veo3_bigrams <- preprocessed_comments %>%
filter(ai_model == "Veo3") %>%
drop_na(cleaned_text) %>%
unnest_tokens(bigram, cleaned_text, token = "ngrams", n = 2) %>%
filter(!is.na(bigram))
# Separate bigrams into two words to filter out stop words
Veo3_bigrams_separated <- Veo3_bigrams %>%
separate(bigram, c("word1", "word2"), sep = " ", remove = FALSE)
# Filter out common stop words and purely numeric words
Veo3_bigrams_filtered_common <- Veo3_bigrams_separated %>%
filter(!word1 %in% stop_words$word) %>%
filter(!word2 %in% stop_words$word) %>%
filter(!str_detect(word1, "^[0-9]+$")) %>%
filter(!str_detect(word2, "^[0-9]+$"))
# Re-unite the filtered individual words back into bigrams
Veo3_bigrams_united_common <- Veo3_bigrams_filtered_common %>%
unite(bigram, word1, word2, sep = " ")
# Count the frequency of each bigram and keep those appearing 4 or more times
Veo3_bigram_counts_common <- Veo3_bigrams_united_common %>%
count(bigram, sort = TRUE) %>%
filter(n >= 4)
# Generate and display a word cloud for Veo3's bigrams
wordcloud(Veo3_bigram_counts_common$bigram, Veo3_bigram_counts_common$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 1))
mtext("Veo3 Bigram", side = 3, line = 0, adj = 0.5, cex = 3.2)
Fig.4-1 Bigram Word Cloud of Sora and Veo3 YouTube Comments: Identifying Key Bigrams after Common Stop Word Removal
# Process Bigrams for Sora Model
# Filter out custom stop words and purely numeric words
Sora_bigrams_filtered_custom <- Sora_bigrams_separated %>%
filter(!word1 %in% custom2_stop_words$word) %>%
filter(!word2 %in% custom2_stop_words$word) %>%
filter(!str_detect(word1, "^[0-9]+$")) %>%
filter(!str_detect(word2, "^[0-9]+$"))
# Re-unite the filtered words back into bigrams
Sora_bigrams_united_custom <- Sora_bigrams_filtered_custom %>%
unite(bigram, word1, word2, sep = " ")
# Count the frequency of each bigram and keep only those appearing 3 or more times
Sora_bigram_counts_custom <- Sora_bigrams_united_custom %>%
count(bigram, sort = TRUE) %>%
filter(n >= 3)
# Generate and display a word cloud for Sora's bigrams.
wordcloud(Sora_bigram_counts_custom$bigram, Sora_bigram_counts_custom$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 1.2))
mtext("Sora Bigram", side = 3, line = 0, adj = 0.5, cex = 3.2)
# Process Bigrams for Veo3 Model
# Filter out custom stop words and purely numeric words
Veo3_bigrams_filtered_custom <- Veo3_bigrams_separated %>%
filter(!word1 %in% custom2_stop_words$word) %>%
filter(!word2 %in% custom2_stop_words$word) %>%
filter(!str_detect(word1, "^[0-9]+$")) %>%
filter(!str_detect(word2, "^[0-9]+$"))
# Re-unite the filtered individual words back into bigrams
Veo3_bigrams_united_custom <- Veo3_bigrams_filtered_custom %>%
unite(bigram, word1, word2, sep = " ")
# Count the frequency of each bigram and keep those appearing 4 or more times
Veo3_bigram_counts_custom <- Veo3_bigrams_united_custom %>%
count(bigram, sort = TRUE) %>%
filter(n >= 4)
# Generate and display a word cloud for Veo3's bigrams
wordcloud(Veo3_bigram_counts_custom$bigram, Veo3_bigram_counts_custom$n,
max.words = 100, colors = brewer.pal(8, "Dark2"), random.order = FALSE,
scale = c(4, 0.7))
mtext("Veo3 Bigram", side = 3, line = 0, adj = 0.5, cex = 3.2)
Fig.4-2 Bigram Word Cloud of Sora and Veo3 YouTube Comments: Identifying Key Bigrams after Custom Stop Word Removal
As revealed by the Sora Bigram word cloud on the left side of Fig. 4-2, key topics observed include: voice mode / voice acting (absence/insufficiency of voice features, anticipation of future integration/development), real people / real time / real life (Sora videos’ high “realism”), uncanny valley (discomfort with video realism; a concrete concept linking to previously inferred negative emotions like ‘scary’ or ‘bad’ from the single-word cloud), social media / film industry / movie industry (significant interest and concern for changes in the existing content industry, seen alongside ‘stock footage’ and ‘film industry’), human creativity / human labor (AI’s impact on human creativity/labor, a more advanced form than ‘jobs’ and ‘artists’ in the single-word cloud), and mind blowing / game changing (strong praise for Sora’s technological level and its potential as an industry-wide ‘game changer’). Expectations for Sora’s potential, vaguely expressed as ‘dream’, ‘imagine’ and ‘future’ in the single-word cloud, materialized into concrete expressions like ‘mind blowing’ and ‘game changing’ in the bigrams, highlighting a strong perception of technological impact and innovation. Moreover, subtle negative emotions related to AI’s realism, such as ‘uncanny valley,’ were captured.
Looking at the Veo3 Bigram word cloud on the right side of Fig. 4-2, major topics inferred include: ai slop / soulless slop (the strongest negative criticism for Veo3 videos, signifying an ‘soulless and meaningless’. low-quality perception of AI-generated content; a concretization of criticisms like ‘slop,’ ‘cooked,’ and ‘fake’ from the single-word cloud), real people / real human / real life (pursuit of realism like ‘real people’; with ‘ai slop’ it questions/criticizes AI’s ability to genuinely mimic reality), bad acting / bad actors (specific criticism of unnaturalness and quality degradation in ‘acting’ or ‘actors’ within AI-generated videos), black mirror / sentinel island (ethical and social concerns about AI technology leading to dystopian futures like ‘Black Mirror’ or uncontrollable dangerous situations like ‘Sentinel Island’; demonstrating a much more specific and serious social implication than the ‘scary’ or ‘bad’ inferred from the single-word cloud), processed food / crunchy spaghetti (criticism implying an artificial, processed, alien, or unnatural feel), content creators / youtube channel (impact on existing content creators and YouTube channels), and human creativity / human art (threat or change to human creativity and art). Veo3’s bigrams appear to convey much more specific and intense criticism compared to Sora. Notably, powerful negative terms like “AI slop” as well as keywords with social, ethical, and dystopian implications such as ‘Black Mirror’ and ‘Sentinel Island’ emerged, indicating deep concerns about AI technology’s side effects and potential risks. Furthermore, criticisms regarding content quality and lack of artistry were more explicitly revealed through specific metaphors like ‘bad acting’ and ’processed food.
To conduct an in-depth analysis of key topics and implications, semantic network analysis was performed. This included co-occurrence network analysis to identify meaningful associations between distinctive words, and bigram network analysis was also conducted. By applying the Phi coefficient to word pairs identified in the co-occurrence network, a Phi-coefficient network was also analyzed, allowing for a more thorough examination of the main topics and implications present in each model’s comments. This network analysis facilitated easy visual tracking of changes through adjustments to stopword inclusion/exclusion, word pair filtering ranges, and Phi coefficient thresholds, enabling effective derivation of key topics and implications through various simulations.
# Sora Co-occurrence Network Analysis
# Prepare Sora comments: filter by model and remove NA texts
Sora_comments <- preprocessed_comments %>%
filter(ai_model == "Sora") %>%
drop_na(cleaned_text)
# Tokenize words, filter stop words and numeric words for Sora comments
Sora_words_filtered <- Sora_comments %>%
unnest_tokens(word, cleaned_text) %>%
filter(!word %in% custom3_stop_words$word) %>%
filter(!str_detect(word, "^[0-9]+$"))
# Calculate pairwise co-occurrence counts for filtered Sora words
# Filters words with a minimum frequency (n >= 7) before calculating pairs
Sora_words_cooc <- Sora_words_filtered %>%
add_count(word) %>%
filter(n >= 7) %>%
pairwise_count(item = word, feature = comments_code, sort = TRUE)
# Create a graph object for Sora co-occurrences.
# Filters co-occurrences by count (n >= 4), calculates degree centrality, and detects communities
Sora_words_cooc_graph <- Sora_words_cooc %>%
filter(n >= 4) %>%
as_tbl_graph(directed = FALSE) %>%
mutate(centrality = centrality_degree(),
group = as.factor(group_infomap()))
set.seed(1234)
# Generate the ggraph plot for Sora's co-occurrence network
ggraph(Sora_words_cooc_graph, layout = "fr") +
geom_edge_link(aes(edge_alpha = n, edge_width = n), color = "gray50", show.legend = FALSE) +
scale_edge_width(range = c(1, 5)) +
geom_node_point(aes(size = centrality, color = group), show.legend = FALSE) +
scale_size(range = c(3, 10)) +
geom_node_text(aes(label = name, size = centrality), repel = TRUE, max.overlaps = 20, show.legend = FALSE) +
scale_size_continuous(range = c(7, 10)) +
theme_graph() +
theme(
plot.background = element_rect(fill = "white", colour = NA),
panel.background = element_rect(fill = "white", colour = NA),
plot.title = element_text(size = 25, hjust = 0.5, face = "bold", family = "sans")) +
labs(title = "Sora Co-occurence Network")
# Veo3 Co-occurrence Network Analysis
# Prepare Veo3 comments: filter by model and remove NA texts
Veo3_comments <- preprocessed_comments %>%
filter(ai_model == "Veo3") %>%
drop_na(cleaned_text)
# Tokenize words, filter stop words and numeric words for Veo3 comments
Veo3_words_filtered <- Veo3_comments %>%
unnest_tokens(word, cleaned_text) %>%
filter(!word %in% custom3_stop_words$word) %>%
filter(!str_detect(word, "^[0-9]+$"))
# Calculate pairwise co-occurrence counts for filtered Veo3 words
# Filters words with a minimum frequency (n >= 7) before calculating pairs
Veo3_words_cooc <- Veo3_words_filtered %>%
add_count(word) %>%
filter(n >= 7) %>%
pairwise_count(item = word, feature = comments_code, sort = TRUE)
# Create a graph object for Veo3 co-occurrences
# Filters co-occurrences by count (n >= 4), calculates degree centrality, and detects communities
Veo3_words_cooc_graph <- Veo3_words_cooc %>%
filter(n >= 4) %>%
as_tbl_graph(directed = FALSE) %>%
mutate(centrality = centrality_degree(),
group = as.factor(group_infomap()))
set.seed(5678)
# Generate the ggraph plot for Veo3's co-occurrence network
ggraph(Veo3_words_cooc_graph, layout = "fr") +
geom_edge_link(aes(edge_alpha = n, edge_width = n), color = "gray50", show.legend = FALSE) +
scale_edge_width(range = c(1, 5)) +
geom_node_point(aes(size = centrality, color = group), show.legend = FALSE) +
scale_size(range = c(3, 10)) +
geom_node_text(aes(label = name, size = centrality), repel = TRUE, max.overlaps = 20, show.legend = FALSE) +
scale_size_continuous(range = c(7, 10)) +
theme_graph() +
theme(
plot.background = element_rect(fill = "white", colour = NA),
panel.background = element_rect(fill = "white", colour = NA),
plot.title = element_text(size = 25, hjust = 0.5, face = "bold", family = "sans")) +
labs(title = "Veo3 Co-occurence Network")
Fig.5 Co-occurrence Network Analysis of Sora and Veo3 YouTube Comments: Identifying Key Themes and Relationships
Without custom stopword processing, words like ‘AI,’ ‘people,’ and ‘video’ generated excessive node connections, hindering analysis. Therefore, words were added or excluded through various simulations. Additionally, by setting limits on the total frequency of individual words and the co-occurrence frequency of word pairs, the visual analysis of the co-occurrence network was facilitated.
Fig. 5’s co-occurrence network analysis clearly reveals word associations and semantic clusters in Sora and Veo3 comments, deeply exposing user perceptions of each AI video model.
Based on the Sora co-occurrence network analysis, users expressed both technological wonder and optimism for the future regarding this AI video model, while also showing anxiety about jobs and social control. Technical terms like ‘AI’, ‘generation’, and ‘video’ connected with ‘future’, ‘imagine’ and ‘dream’, indicating high expectations for Sora’s innovation and its future potential. Its links to ‘artists’, ‘industry’ and ‘film’ showed interest in its positive impact on creative fields. However, ‘jobs’ and ‘scary’ connected to ‘human’ and ‘tech’ revealed concerns about job replacement and general societal unease due to technological advancement. ‘Misinformation’, ‘control’ and ‘automation’ also suggested potential worries about misuse and social changes. Furthermore, ‘voice’ and ‘mode’ clusters indicated expectations for improved audio features, and mentions of body parts like ‘hands’ and ‘fingers’ reflected demands for better realism in details. Ultimately, user perception of Sora complexly combines optimism about technology’s potential with a deep consideration of the social and ethical changes AI might bring.
Veo3 co-occurrence network analysis reveals users hold a distinctly critical view of this AI video model. Strong connections among terms like ‘ai slop’, ‘soulless’, ‘cooked’, ‘fake’ and ‘bad acting’ predominantly show Veo3 content is perceived as soulless, fake, and low quality. Metaphors such as ‘processed food’ or ‘crunchy spaghetti’ further emphasize dissatisfaction with its artificial and unnatural feel. Links among ‘black mirror’, ‘sentinel island’, ‘scary’ and ‘bad’ express deep ethical concerns and fears about dystopian social changes or uncontrolled risks that Veo3’s technology might bring. Furthermore, the connection of ‘real people’ and ‘real human’ with negative keywords reflects user assessment that Veo3 struggles to convincingly mimic reality, resulting in unnatural outputs. Terms like ‘film industry’, ‘hollywood’, ‘money’ and ‘jobs’ appearing with negative keywords, suggest a strong skeptical view of AI’s impact on these industries. Additionally, the ‘google’ node indicates user awareness of the developer. Ultimately, Veo3’s perception is dominated by strong criticism of its content’s artistry and lack of meaning rather than its technical perfection, combined with complex and specific concerns about the social and ethical risks of AI technology.
# Sora Bigram Network Analysis
# Process Sora bigrams: tokenize, separate, filter stop words and numeric words
Sora_bigrams_processed <- Sora_comments %>%
unnest_tokens(bigram, cleaned_text, token = "ngrams", n = 2) %>%
filter(!is.na(bigram)) %>%
separate(bigram, c("word1", "word2"), sep = " ", remove = FALSE) %>%
filter(!word1 %in% custom3_stop_words$word) %>%
filter(!word2 %in% custom3_stop_words$word) %>%
filter(!str_detect(word1, "^[0-9]+$")) %>%
filter(!str_detect(word2, "^[0-9]+$"))
# Count co-occurrences of filtered bigrams for Sora
Sora_bigram_cooc <- Sora_bigrams_processed %>%
count(word1, word2, sort = T) %>%
na.omit()
# Create a graph object for Sora bigrams, filter by frequency, and calculate network metrics
Sora_bigram_graph <- Sora_bigram_cooc %>%
filter(n >= 3) %>%
as_tbl_graph(directed = FALSE) %>%
mutate(centrality = centrality_degree(),
group = as.factor(group_infomap()))
set.seed(1234)
a <- grid::arrow(type = "closed", length = unit(.15, "inches"))
# Generate the ggraph plot for Sora's bigram network
ggraph(Sora_bigram_graph, layout = "fr") +
geom_edge_link(aes(edge_alpha = n, width = n), show.legend = FALSE,
arrow = a, end_cap = circle(.07, 'inches')) +
scale_edge_width_continuous(range = c(1, 5)) +
geom_node_point(aes(size = centrality, color = group), show.legend = FALSE) +
scale_size(range = c(3, 10)) +
geom_node_text(aes(label = name, size = centrality), repel = TRUE, max.overlaps = 20, show.legend = FALSE) +
scale_size_continuous(range = c(7, 10)) +
theme_graph() +
theme(
plot.background = element_rect(fill = "white", colour = NA),
panel.background = element_rect(fill = "white", colour = NA),
plot.title = element_text(size = 25, hjust = 0.5, face = "bold", family = "sans")) +
labs(title = "Sora Bigram Network")
# Veo3 Bigram Network Analysis
# Process Veo3 bigrams: tokenize, separate, filter stop words and numeric words.
Veo3_bigrams_processed <- Veo3_comments %>%
unnest_tokens(bigram, cleaned_text, token = "ngrams", n = 2) %>%
filter(!is.na(bigram)) %>%
separate(bigram, c("word1", "word2"), sep = " ", remove = FALSE) %>%
filter(!word1 %in% custom3_stop_words$word) %>%
filter(!word2 %in% custom3_stop_words$word) %>%
filter(!str_detect(word1, "^[0-9]+$")) %>%
filter(!str_detect(word2, "^[0-9]+$"))
# Count co-occurrences of filtered bigrams for Veo3
Veo3_bigram_cooc <- Veo3_bigrams_processed %>%
count(word1, word2, sort = T) %>%
na.omit()
# Create a graph object for Veo3 bigrams, filter by frequency, and calculate network metrics
Veo3_bigram_graph <- Veo3_bigram_cooc %>%
filter(n >= 3) %>%
as_tbl_graph(directed = FALSE) %>%
mutate(centrality = centrality_degree(),
group = as.factor(group_infomap()))
set.seed(5678)
a <- grid::arrow(type = "closed", length = unit(.15, "inches"))
# Generate the ggraph plot for Veo3's bigram network
ggraph(Veo3_bigram_graph, layout = "fr") +
geom_edge_link(aes(edge_alpha = n, width = n), show.legend = FALSE,
arrow = a, end_cap = circle(.07, 'inches')) +
scale_edge_width_continuous(range = c(1, 5)) +
geom_node_point(aes(size = centrality, color = group), show.legend = FALSE) +
scale_size(range = c(3, 10)) +
geom_node_text(aes(label = name, size = centrality), repel = TRUE, max.overlaps = 20, show.legend = FALSE) +
scale_size_continuous(range = c(7, 10)) +
theme_graph() +
theme(
plot.background = element_rect(fill = "white", colour = NA),
panel.background = element_rect(fill = "white", colour = NA),
plot.title = element_text(size = 25, hjust = 0.5, face = "bold", family = "sans")) +
labs(title = "Veo3 Bigram Network")
Fig.6 Bigram Network Analysis of Sora and Veo3 YouTube Comments: Identifying Key Themes and Relationships
The Sora bigram network, shown on the left side of Fig.6, prominently features positive technological marvel and future-oriented discussions. The core concept of ‘AI generated’ forms strong connections at the network’s center with ‘content’ and ‘creation’ indicating high interest in Sora’s generation technology itself. Notably, ‘voice mode’ links with ‘voice’ showing users’ high expectations for audio capabilities despite visual superiority and active discussion on future integration possibilities. Users’ astonishment at the overwhelming ‘realism’ of Sora’s videos is evident in related words, but this realism also comes with a subtle discomfort like ‘uncanny valley’. Terms such as ‘film industry’, ‘industry’ and ‘stock footage’ illustrate concrete discussions about Sora’s profound impact and the changes it will bring to existing content industries. Furthermore, deep contemplation on AI’s impact on human creativity and labor is clearly evident in the bigrams, revealed through connections involving ‘industry’, ‘money’ and ‘person’. Expressions like ‘mind blowing’ reflect a strong recognition of Sora’s technological level and its status as a revolutionary force poised to transform entire industries. Overall, while Sora generates high expectations for its technological potential and the future, concerns about social impacts like jobs are also actively discussed.
The Veo3 bigram network, shown on the right side of Fig.6, primarily revolves around intense criticism of technological limitations and content quality, alongside significant social/ethical concerns. Most notably, ‘ai slop’ forms a strong cluster with ‘soulless slop’, ‘cooked’ and ‘fake’ revealing a very direct and strong critical perception that AI-generated content is ‘soulless and meaningless’ of low quality. This connects with concrete metaphors like ‘bad acting’, ‘processed food’ and ‘crunchy spaghetti’ clearly expressing dissatisfaction with unnaturalness and quality degradation in video content. Keywords such as ‘black mirror’, ‘sentinel island’ and ‘interdimensional cable’ group together with ‘scary’ and ‘bad’, suggesting a dominant presence of deep and specific ethical and social concerns that Veo3’s technological advancement could lead to a ‘dystopian future’ or ‘uncontrollable risks’. The association of ‘real people’ and ‘real human’ with terms like ‘ai slop’ or ‘bad acting’ reinforces user evaluations that Veo3 has limitations in mimicking real humans, resulting in unnatural or ‘fake’ outputs. Industry-related terms like ‘film industry’, ‘hollywood’, ‘money’ and ‘jobs’ are still mentioned, but their appearance alongside negative keywords like ‘ai slop’ indicates a strong skeptical view regarding AI’s impact on existing industries and employment. In summary, while comments on Veo3 do mention its technological completeness (especially audio), strong criticism of content’s ‘artistry’ or ‘lack of meaning’ (“AI slop”) is prominent, and concerns about social impacts like ‘jobs’ are phenomena common to both models.
# Sora Phi Coefficient Network Analysis
# Calculate pairwise phi correlations for words in Sora comments
# Filters words with frequency >= 6 and keeps correlations > 0.3
Sora_words_phi <- Sora_words_filtered %>%
add_count(word) %>%
filter(n >= 7) %>%
pairwise_cor(item = word, feature = comments_code, sort = TRUE) %>%
filter(correlation > 0.27)
# Create a graph object for Sora based on phi coefficients
# Calculate degree centrality and perform community detection
Sora_graph_phi <- Sora_words_phi %>%
as_tbl_graph(directed = FALSE) %>%
mutate(centrality = centrality_degree(),
group = as.factor(group_infomap()))
set.seed(1234)
# Generate the ggraph plot for Sora's phi coefficient network
ggraph(Sora_graph_phi, layout = "fr") +
geom_edge_link(aes(edge_alpha = correlation, edge_width = correlation), color = "gray50", show.legend = FALSE) +
scale_edge_width(range = c(1, 5)) +
geom_node_point(aes(size = centrality, color = group), show.legend = FALSE) +
scale_size(range = c(3, 10)) +
geom_node_text(aes(label = name, size = centrality), repel = TRUE, max.overlaps = 20, show.legend = FALSE) +
scale_size_continuous(range = c(7, 10)) +
theme_graph() +
theme(
plot.background = element_rect(fill = "white", colour = NA),
panel.background = element_rect(fill = "white", colour = NA),
plot.title = element_text(size = 25, hjust = 0.5, face = "bold", family = "sans")) +
labs(title = "Sora Phi-coefficient Network")
# Veo3 Phi Coefficient Network Analysis
# Calculate pairwise phi correlations for words in Veo3 comments
# Filters words with frequency >= 6 and keeps correlations > 0.3
Veo3_words_phi <- Veo3_words_filtered %>%
add_count(word) %>%
filter(n >= 7) %>%
pairwise_cor(item = word, feature = comments_code, sort = TRUE) %>%
filter(correlation > 0.27)
# Create a graph object for Veo3 based on phi coefficients
# Calculate degree centrality and perform community detection
Veo3_graph_phi <- Veo3_words_phi %>%
as_tbl_graph(directed = FALSE) %>%
mutate(centrality = centrality_degree(),
group = as.factor(group_infomap()))
set.seed(5678)
# Generate the ggraph plot for Veo3's phi coefficient network
ggraph(Veo3_graph_phi, layout = "fr") +
geom_edge_link(aes(edge_alpha = correlation, edge_width = correlation), color = "gray50", show.legend = FALSE) +
scale_edge_width(range = c(1, 5)) +
geom_node_point(aes(size = centrality, color = group), show.legend = FALSE) +
scale_size(range = c(3, 10)) +
geom_node_text(aes(label = name, size = centrality), repel = TRUE, max.overlaps = 20, show.legend = FALSE) +
scale_size_continuous(range = c(7, 10)) +
theme_graph() +
theme(
plot.background = element_rect(fill = "white", colour = NA),
panel.background = element_rect(fill = "white", colour = NA),
plot.title = element_text(size = 25, hjust = 0.5, face = "bold", family = "sans")) +
labs(title = "Veo3 Phi-coefficient Network")
Fig.7 Phi-coefficient Network Analysis of Sora YouTube Comments: Identifying Key Themes and Relationships
The Phi-coefficient network analysis (Fig. 7) visualizes statistically significant correlations between words in Sora and Veo3 YouTube comments (i.e., words appearing together more often than by chance), thereby more precisely revealing user perceptions of each AI video model.
The Sora Phi-coefficient network reveals a complex interplay of positive expectations regarding technological innovation alongside deep considerations for its potential social and ethical ripple effects. Through this graph, several key word clusters and their relationships can be identified.
Firstly, words like ‘technological’ and ‘progress’ are connected with terms such as ‘vision’, ‘planet’ and ‘movement’, suggesting an anticipation for future technologies and the changes they will bring. Simultaneously, a tightly knit cluster of ‘workers’, ‘family’ and ‘automation’ clearly demonstrates significant interest in and concern about AI’s impact on ‘labor’ and ‘employment’. This concern extends further to ‘political’ and ‘banned’ through their connections, indicating broader societal and political issues, as well as regulatory discussions that automation might trigger. Notably, the ‘harm’, ‘dangerous’ and ‘children’ cluster explicitly highlights ethical concerns that technological development could be harmful or dangerous for children.
Secondly, words related to media content and commercial use also form important clusters. ‘Stock’, ‘footage’ and ‘portfolio’ suggest discussions about the impact of AI video generation technologies like Sora on the ‘stock footage’ and ‘portfolio’ markets within the media industry. Additionally, the ‘blurred’, ‘logo’ and ‘car’ cluster indicates practical issues related to commercial use, such as logo blurring in videos.
Thirdly, emotional responses to AI-generated content are also present. Words like ‘uncanny’, ‘vibes’ and ‘westworld’ suggest an awareness of the “uncanny valley” effect caused by the realism of AI-generated content. This cluster is indirectly linked to ‘harm’ and ‘dangerous’, indicating concerns about potential harm that goes beyond mere discomfort.
In conclusion, the overall perception of Sora, as depicted in this network, includes considerable awe at technological advancement. However, it also significantly correlates with deep reflection and concern about AI’s complex impact on society and humans. Specifically, the network reveals active discussions around the negative ripple effects of technology, changes in the labor market, ethical issues, and the commercial use of content.
In contrast, The Veo3 Phi-coefficient network reveals core correlations driven by strong criticism of content quality and deep concerns about the severe social and ethical risks posed by AI technology. Key word clusters and their meanings, as shown in the graph, are as follows.
Firstly, negative evaluations of content quality are prominent. ‘Soulless’ and ‘slop’ (likely referring to AI slop) are connected with ‘machines’, ‘disgusting’ and ‘effort’, suggesting a very strong and concrete negative assessment that Veo3 content is ‘soulless’, ‘disgusting’ and ‘lacking effort’. This is further linked to ‘truth’, ‘manipulated’ and ‘public’, reflecting concerns about the authenticity of generated content and the potential for public manipulation. The connection of ‘artists’ and ‘hire’ with ‘soulless’ indicates discussions about the lack of artistry in AI-generated content and its relationship with human artists.
Secondly, deep concerns about the social and ethical threats that AI technology might bring are evident. The ‘black’, ‘mirror’, ‘episode’ cluster alludes to dystopian futures, while ‘sentinel’ and ‘north island’ evoke associations with uncontrollable risks or isolated threats. These clusters connect with ‘unnatural’, ‘interdimensional’ and ‘cable’, illustrating anxieties about reality distortion or uncontrollable scenarios stemming from AI technology. Notably, words like ‘uncanny’ and ‘valley’ suggest discomfort caused by the realism of AI-generated content, further expressing confused or negative emotions when paired with exclamations like ‘worry’, ‘crap’ and ‘holy’.
Thirdly, words related to technical aspects also appear. ‘Hardware’, ‘graphics’ and ‘processing’ indicate an interest in the underlying technology of AI models, while ‘data’ and ‘training’ represent elements related to AI model learning. ‘Budget’, ‘low’ and ‘power’ likely reflect discussions about the cost-effectiveness of content generation or certain technical limitations. Additionally, ‘subscription’, ‘credits’ and ‘half’ suggest economic aspects related to service usage.
In summary, user perception of Veo3 clearly shows that strong criticism of the content’s ‘artistry’ or ‘lack of meaning’ is significantly correlated with much deeper and more specific concerns about the social and ethical side effects AI might bring. This emphasizes that users evaluate AI video technology from a complex perspective encompassing ‘creativity’, ‘artistry’, ‘humanity’ and ‘social responsibility’, going beyond simple technical ‘accuracy’.
This study aimed to derive key insights for the future development of AI by deeply analyzing public reactions to new AI video models. The conclusions for each research question are summarized as follows.
① What emotions, reactions, and perceptions do people show regarding the realism of AI-generated videos?: User perceptions of AI video realism are complex. While generally positive, negative emotions increased with higher realism. Sora evoked wonder at technological innovation, but also subtle discomfort, like confusion between reality and fiction. Veo3, despite its high realism, faced criticism for lacking ‘artistry/meaning’ and raised deep concerns about potential ‘Black Mirror’-like risks. This suggests that as realism advances, users evaluate technology more critically.
② What are the key public discussions about AI video models and their outputs, and what factors drive them?: Key public discussions extend beyond technical completeness to include social and ethical implications. Sora discussions centered on positive potential and industry impact. In contrast, Veo3 discussions were dominated by criticism of content quality (‘AI slop’) and profound concerns about potential risks. ‘Jobs’ was a shared core concern for both models, indicating that AI’s social repercussions, beyond just realism, are significant drivers of public discourse.
③ What practical insights and implications for the direction of AI technology development can be derived from these analysis results?: AI video technology development should aim in three directions: First, prioritize ‘artistry, creativity, and human sensibility’ to address issues like the ‘uncanny valley.’ Second, companies must communicate transparently and responsibly about user concerns, including jobs, social control, and ethical issues. Third, AI should ultimately strive for ‘human-centered AI,’ focusing on artistic/human value and societal impact beyond functional accuracy.
In Conclusion, The evolution of AI video technology elicits both wonder and deep concern from users. Technology adoption hinges not merely on realism, but on acknowledging content’s artistic and human value, alongside social responsibility. Therefore, future AI development must integrate creativity, empathy, and ethical considerations to build user trust and manage societal impacts positively, moving towards a ‘human-centered AI’. These insights pose crucial questions not only for technology developers but also for all users consuming AI content and society as a whole observing this trend.
Thank you.