1 million Reddit comments from 40 subreddits. (n.d.). Retrieved April 25, 2025, from https://www.kaggle.com/datasets/smagnan/1-million-reddit-comments-from-40-subreddits
Ali, N. T., Hassan, K. F., Abdullah, M. N., & Al-Hchimy, Z. S. (2024). The Application of Random Forest to the Classification of Fake News. BIO Web of Conferences, 97, 00049. https://doi.org/10.1051/bioconf/20249700049
Allyn, B. (2025, April 7). How a false X post about pausing tariffs led to multitrillion-dollar market swings. NPR. https://www.npr.org/2025/04/07/nx-s1-5355055/tariffs-markets-x-social-media
Antony Vijay, J., Anwar Basha, H., & Arun Nehru, J. (2021). A Dynamic Approach for Detecting the Fake News Using Random Forest Classifier and NLP. In V. Singh, V. K. Asari, S. Kumar, & R. B. Patel (Eds.), Computational Methods and Data Engineering (pp. 331–341). Springer. https://doi.org/10.1007/978-981-15-7907-3_25
Corporation, M., & Weston, S. (2022). doParallel: Foreach Parallel Adaptor for the “parallel” Package (p. 1.0.17) [Dataset]. https://doi.org/10.32614/CRAN.package.doParallel
Del Vicario, M., Bessi, A., Zollo, F., Petroni, F., Scala, A., Caldarelli, G., Stanley, H. E., & Quattrociocchi, W. (2016). The spreading of misinformation online. Proceedings of the National Academy of Sciences, 113(3), 554–559. https://doi.org/10.1073/pnas.1517441113
Dunbar, R. I. M., Arnaboldi, V., Conti, M., & Passarella, A. (2015). The structure of online social networks mirrors those in the offline world.Social Networks, 43, 39–47. https://doi.org/10.1016/j.socnet.2015.04.005
Frith, C. D. (2008). Social cognition. Philosophical Transactions of the Royal Society B: Biological Sciences, 363(1499), 2033–2039. https://doi.org/10.1098/rstb.2008.0005
Gallagher, S. (2020). Interaction. In S. Gallagher (Ed.), Action and Interaction (p. 0). Oxford University Press. https://doi.org/10.1093/oso/9780198846345.003.0006
Gentina, E., Chen, R., & Yang, Z. (2021). Development of theory of mind on online social networks: Evidence from Facebook, Twitter, Instagram, and Snapchat. Journal of Business Research, 124, 652–666. https://doi.org/10.1016/j.jbusres.2020.03.001
Jockers, M. (2015). syuzhet: Extracts Sentiment and Sentiment-Derived Plot Arcs from Text (p. 1.0.7) [Dataset]. https://doi.org/10.32614/CRAN.package.syuzhet
Kanai, R., Bahrami, B., Roylance, R., & Rees, G. (2011). Online social network size is reflected in human brain structure. Proceedings of the Royal Society B: Biological Sciences, 279(1732), 1327–1334. https://doi.org/10.1098/rspb.2011.1959
Kuhn, M. (2008). Building Predictive Models in R Using the caret Package. Journal of Statistical Software, 28, 1–26. https://doi.org/10.18637/jss.v028.i05
Liaw, A., & Wiener, M. (2002). Classification and Regression by randomForest.
Media Bias Chart. (n.d.). Retrieved April 27, 2025, from https://app.adfontesmedia.com/chart/interactive?utm_source=adfontesmedia&utm_medium=website
Meshi, D., Tamir, D. I., & Heekeren, H. R. (2015). The Emerging Neuroscience of Social Media. Trends in Cognitive Sciences, 19(12), 771–782. https://doi.org/10.1016/j.tics.2015.09.004
Mohammad, S. M., & Turney, P. D. (2013). Crowdsourcing a Word-Emotion Association Lexicon (No. arXiv:1308.6297). arXiv. https://doi.org/10.48550/arXiv.1308.6297
Park, C., Majeed, A., Gill, H., Tamura, J., Ho, R. C., Mansur, R. B., Nasri, F., Lee, Y., Rosenblat, J. D., Wong, E., & McIntyre, R. S. (2020). The Effect of Loneliness on Distinct Health Outcomes: A Comprehensive Review and Meta-Analysis. Psychiatry Research, 294, 113514. https://doi.org/10.1016/j.psychres.2020.113514
Premack, D., & Woodruff, G. (1978). Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4), 515–526. https://doi.org/10.1017/S0140525X00076512
Ramos, J. (2003). Using TF-IDF to determine word relevance in document queries. ResearchGate. https://www.researchgate.net/publication/228818851_Using_TF-IDF_to_determine_word_relevance_in_document_queries
R/The_Donald. (2025). In Wikipedia. https://en.wikipedia.org/w/index.php?title=R/The_Donald&oldid=1287208539
Sapolsky, R. (2017). Behave: The biology of humans at our best and worst. Penguin Press.
Selivanov, D., Bickel, M., & Wang, Q. (2023). text2vec: Modern Text Mining Framework for R (p. 0.6.4) [Dataset]. https://doi.org/10.32614/CRAN.package.text2vec
Silge, J., & Robinson, D. (2016). tidytext: Text Mining and Analysis Using Tidy Data Principles in R. The Journal of Open Source Software, 1(3), 37. https://doi.org/10.21105/joss.00037
Swire-Thompson, B., & Lazer, D. (2020). Public Health and Online Misinformation: Challenges and Recommendations. Annual Review of Public Health, 41(1), 433–451. https://doi.org/10.1146/annurev-publhealth-040119-094127
Wickham, H. (2016). ggplot2: Create Elegant Data Visualisations Using the Grammar of Graphics (Version 3.5.2) [Computer software]. https://cran.r-project.org/web/packages/ggplot2/index.html
Wickham, H., Wickham H, Averick M, Bryan J, Chang W, McGowan LD, François R, Grolemund G, Hayes A, Henry L, Hester J, Kuhn M, Pedersen TL, Miller E, Bache SM, Müller K, Ooms J, Robinson D, & Seidel DP, Spinu V, Takahashi K, Vaughan D, Wilke C, Woo K, Yutani H. (2019). tidyverse: Easily Install and Load the “Tidyverse” (Version 2.0.0) [Computer software]. https://cran.r-project.org/web/packages/tidyverse/index.html
A link to my presentation slides can be found here: Final Presentation Slides
But then I was like all these other people to understand this had to know that I’m a senior. This is my last meet ever, and I’ve had a really positive experience in diving, and it’s not like they took the time to sit there and think. Hmm! Why did Abby post this like? Let me think about this, but they’re able to do it instantaneously while scrolling on their phone. And I was like, Wow, that’s
kind of crazy. I don’t know. That might just be me. But so one thing that I thought about, so we’ll going back. Now, we can talk a little bit about theory. So social cognition studies how humans interact in social groups. So most of this goes on in person. And one of my one of the sub categories in social cognition that interests me particularly is
is theory of mind which explains the ability to understand the thoughts and beliefs of others. So I’m particularly interested in this because it just sort of explains our everyday interactions like you kind of move through every day being able to think, oh, like this person is doing this because x or do you feel y? Because this happened and
good. Okay. And so one of the modern understandings of theory of mind comes from an embodied approach. So it’s that we have a body and a mind, and we’re operating in this world and sort of our experience in this world helps us understand other people. So instead of like simulating or theorizing about what’s going on in other people’s minds. We understand it based on our understanding of bodies in the world. But then, with social media, you don’t have a body in the world, in social media. It’s a digital and asynchronous place. So
I hypothesize that the the action of theory of mind doesn’t occur the same way
online as it does in person. And one way I thought to look at that was with social media data.
So I found this recent review that talked about the 5 ways that social cognition is used in social media. So the 5 ways include broadcasting information. So that’s like posting, receiving feedback through comments, observing others, providing feedback to others, like commenting on their posts and then comparing, which is like self referencing to others.
And so this gave me like this. We have all this social media data because people are on it all the time posting all the time. So I was like, this seems like a great source of data to learn about social cognition in the digital environment.
So social media data has been used a lot in a ton of like computer science and cognitive science studies, graphing social networks, and sort of understanding how information flows and how people understand others. So I thought that this is not like a totally new field. But it’s something that’s been well established and something that we can use to look at social cognition.
So then, my big question was, how does social cognition in person. Compare to social cognition on the Internet with a specific focus in theory of mind, and how we understand the mental states of others when it’s an asynchronous digital environment as opposed to in person, where we have an embodied approach.
So I was thinking about ways that I could do this on a small scale. And I was like, Oh, I’ll create a misinformation classifier, because that’s something that sounds
like relevant and sort of something that I could manageable for me, and so, while I was doing it, I was scrolling on and looking at the news, and I found this post. That was this article is how false X post about pausing tariffs led to multi trillion dollar market swings. And I was thinking about this in terms of misinformation and theory of mind. I’m like to understand this information. You have to understand, like the source of your information, the intentions of that source so like. For this example, the post came from
like a source that wasn’t. It was verified, but in the sense of like ex verified. But it wasn’t like a major news outlet or anything, so the people who saw it would have to decide whether this source is like worthy, or if it could be misinformation which clearly they did not do such a good job on.
But so I found a kaggle data set a reddit data set on Kaggle that had a million reddit comments from 40 different subreddits. And so it had 4 different columns that had the subreddit
the body of the text, the score, which is the difference of the Upvotes and Downvotes, and then the controversiality score, which is just a 1 or 0, based on whether it was controversial or not, and the controversiality was determined whether the scores had a similar number of Upvotes or down votes, making it controversial or not.
But before I get into my misinformation classifier, I would like to make many disclaimers. So first, st this data set wasn’t tagged for misinformation classification. So I decided to like hard code it. So this is very much a prototype, and you’ll see, as I go through many reasons as to why that is. But so this is. Not only are the comments isolated from the context, but I also hard coded them for misinformation. So this model is not very applicable to anything. It was just sort of like a tasting of what we could do with social media data.
So I went through and did some Eda and determined that the most common score was 2 sort of like in the middle. So there’s like a lot ones that had downloads and ones that had
outboats.
I don’t know. It was just, and then I looked at controversiality, and I found that 97% of comments are not tagged as controversial, and with the average score around 0 which makes sense. If you’re looking at comments at a similar number of votes and downvotes.
I did sentiment, analysis as well, and the most common sentiments are positive, negative, and trust, and I found it interesting. That trust was the 3rd most common sentiment, considering I was going to use this to look at misinformation.
So I use a random Forest Misinformation Classifier. I used Random Forest because I found it be good to like average multiple decision trees to sort of have a better.
just a better model in general. I used Claude to help me code this, and it used parallel processing to speed up training on my computer. And then it also used tf, idf, which I’d never heard of before until this, which is stands for term frequency, inverse document frequency. And so because, like, I was initially learned to like, look at taking out stop words when doing language stuff so. But tf, I found really interesting because it uses the frequency of it in the comments, and then, like in all the text sort of seems like a more balanced and better way to do it than using the hard coded stop words
and then, as I mentioned before, I hard coded the misinformation, because I’m not exactly sure in retrospect why, I thought this was a good idea, but we hard coded words, such as conspiracy, hoax, fake, scam, and then also flagged some
subreddits that might be likely to have higher misinformation levels, such as the conspiracy subreddit politics and world news. And then I, using both the words we we set and the subreddits we called, told to assign different probabilities of misinformation, based on the information. So like, for example, if it had the potential misinformation words and came from that.
So I’ve read it. Then it was given an 80% chance of being misinformation. And then it was for other ones as well. And then I decided to use an 80 20 train test split.
So the results. It claims to have 96% accuracy with precision of one and a recall of 85.7. But that is, of course, all because it does exactly what I told it to do where I trained it on. If it has this word in it, then it’s misinformation. If it doesn’t have this word, then it’s not so. It’s sort of explained by my testing. So we use a prediction. So I put in, it is a conspiracy theory. And that’s 99.5% potential misinformation, which, because it has the keyword in it that I told it was misinformation.
But then I put in the earth is flat, which we all know isn’t true, but it thinks it’s just as factual as the same statement. This is a factual information.
So this model, again very much prototype, but it doesn’t have very many applications. But I was sort of thinking of. This is one thing that, like social media data, could be used for
so limitations. I’ve mentioned already. Comments are isolated from the context. You can’t really understand the interactions that are going on between people. And then the misinformation. Comments were hard coded. But then, so I wanted to talk a little bit about using social media data to study not only misinformation, but theory of mind in general, but from social media, because it’s an asynchronous digital place. The mental states of users are inaccessible to researchers. So one sort of
way I thought to go around about this was if you could find volunteers who are willing to offer their social media posts and data, but with like annotation so like, I would submit my Instagram post and be like a post commemorating the end of my diving career. And so they could. Researchers could potentially use that in looking at social media data, to understand like this is the information that was on the web. This is the information that, like the motivation, or like the thought process behind it, and sort of understand how people interact using this without having that background information.
And yeah, so these are some of my sources. Take questions, grab the recording correctly.
Yes, or don’t really comment. You know what ways that they can be like, how they can test the theory of mind. Could they test it based on their like Instagram feed or something like that? Yeah, I think it’d be really interesting, because you’re right. There are a lot of people who are just observers and aren’t like broadcasting or commenting. And so I think it’d be more interesting to like come up with like
like a survey wouldn’t be. I don’t feel like would be ideal, but I feel like it’d be really interesting to hear about like those people and see what they’re gaining from it. So maybe a survey. But that’s my thought.
Did you consider it like that last point of like having people like submit like their own posts and like, think, oh, this is why like the idea behind it? Wouldn’t that also, still.
like, continue like the misinformation spread of people like have a misunderstanding already, like a like a like. For example, if I generally.
I don’t believe the earth is flat, and I submit a flat earth post. And I say, Oh, the earth is flat like this is real like, wouldn’t that then, further exemplify like those biases? So that’s a really interesting point.
Who’s next? Go ahead. I was thinking about for like for like studying it, it wouldn’t be sort of like. So like you would submit like, Oh, this is a fire so like researchers would not understand like that. That’s your true belief. And then, like more, look at the interactions of like other people interacting with the post. So it’s like someone else looks in like, oh, he’s a valuable source. The earth must be flat, or it’s like, Oh, this is like a source that might not be so reliable. I should take this
grain of salt.
Charlie.
it looked like you had some like sentiment analysis data. Do you ever like consider like posts that are more emotional like that are evoking anger or things like that might be more likely to have misinformation just because they’re trying to listen to that like emotional reaction. Yeah, that’s definitely an excellent point and excellent thought, and would make a lot of sense. But I’m not the best at coding, and would not know how to like explore, especially with my data set being not really used for what I used it for as well.
An awful lot of science has been using other people’s data set for questions that they never asked thanks.
So one quick comments in cognition and computation, a huge field. You were involved for certain kinds of situations you’re involved. For here and now you were involved to interact with other people.
Socially.
you’re involved to figure out what’s going on in their minds. Right? But you’re now in situations that you’re not evolved for where you don’t just have joint attention and joint attention. People in the scene but blended classic joint attention where some of those people are 200 years old
in the sense that you’re reading their writings. They’re not here, they’re not. Now. You have some people who are broadcasters from Paris. They’re not here. They’re not. You live in a world where your scenes of joint attention are not the ones usually that you were evolved for, and the question of how
your evolved cognition is going to work when you are in scenes involving other agents that are not the ones that you were evolved to interact with, not in the positions, not in the time, not in the place. Maybe they’re not even human.
This is this major question in cognition and computation. Take it away.
# Misinformation Detection Model for Reddit Comments Using Random Forests
# Using the Kaggle dataset: 1 million Reddit comments from 40 subreddits
# load required packages
library(tidyverse)
library(tidytext)
library(text2vec)
library(tm)
library(syuzhet)
library(kableExtra)
library(randomForest)
library(caret)
library(readr)
library(doParallel)
library(ggplot2)
library(plotly)
library(DT)
library(heatmaply)
library(wordcloud2)
library(ROCR)
library(viridis)
library(htmlwidgets)
library(corrplot)
library(igraph)
library(ggraph)
# read in data
reddit_data <- read.csv("kaggle_RC_2019-05.csv")
# make controversiality a factor
reddit_data <- reddit_data %>% mutate(controversiality = as.factor(controversiality))
# look at dataset
str(reddit_data)
## 'data.frame': 1000000 obs. of 4 variables:
## $ subreddit : chr "gameofthrones" "aww" "gaming" "news" ...
## $ body : chr "Your submission has been automatically removed because all post titles must begin with one hard-bracketed spoil"| __truncated__ "Dont squeeze her with you massive hand, you mean giant." "It's pretty well known and it was a paid product placement. Hamilton advertised the watch around the movie and "| __truncated__ "You know we have laws against that currently correct? Or are you just willfully ignorant of gun laws in the US." ...
## $ controversiality: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
## $ score : int 1 19 3 10 1 2 7 9 3 3 ...
# plot of comment score frequency
ggplot(reddit_data, aes(x = score)) +
geom_histogram(binwidth = 50, fill = "tomato2") +
xlim(-500, 1000) +
ylim(0, 10000) +
labs(title = "Score Histogram", xlab = "Score", ylab = "Count") +
theme_minimal()
summary(reddit_data$score)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## -889.00 1.00 2.00 11.51 4.00 35619.00
# most common score
most_common_scores <- data.frame(table(reddit_data$score)) %>% arrange(desc(Freq))
head(most_common_scores, 10)
## Var1 Freq
## 1 1 386651
## 2 2 163089
## 3 3 83853
## 4 0 43228
## 5 4 36744
## 6 5 34012
## 7 6 25164
## 8 7 19463
## 9 8 15306
## 10 -1 14688
# comment from highest score entry
# reddit_data[which.max(reddit_data$score),2]
# comment from lowest score entry
# reddit_data[which.min(reddit_data$score),2]
# plot of controversiality scores
ggplot(reddit_data, aes(x = controversiality)) +
geom_bar(fill = c("lightblue", "salmon")) +
labs(title = "Comment Controversiality",
xlab = "Controversiality",
ylab = "Count") +
theme_minimal()
table(reddit_data$controversiality)
##
## 0 1
## 970417 29583
# looking at controversial comments
controversial_comments <- reddit_data %>% filter(controversiality == 1)
summary(controversial_comments$score)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## -207.0000 -1.0000 0.0000 0.5027 2.0000 250.0000
# controversial comment distribution
ggplot(controversial_comments) +
geom_histogram(aes(x = score), binwidth = 25, fill = "orange1" ) +
theme_minimal() +
labs(title = "Controversial Comment Scores")
controversial <- which(reddit_data$controversiality == 1)
# examples of controversial comments
reddit_data[controversial[1:5],2]
## [1] "Can I ask what more you expect out of someone giving an interview? You're there for whoever the guest is not the host."
## [2] "So you’d say 26 year old Buddy Hield isn’t a part of the Kings young core?"
## [3] "Not in my findings"
## [4] "A competent writer could of finished the series in okay manner at least, some of D&Ds choices are so bad they seem deliberate."
## [5] "> If you remove suicides and gang violence our \"gun violence epidemic\" suddenly disappears entirely.\n\nMass-shootings in Universities are not by gangs, nor suicide cases (despite shooters routinely taking their own lives afterwards). It is something that just doesn't happen on the scale it does in America, and a large factor in that is the ease of gun access.\n\n6 people a year (wherever you pulled that stat) is 6 too many."
# convert comments to all lowercase, remove punctuation
reddit_data$body <- tolower(reddit_data$body)
reddit_data$body <- gsub("[[:punct:]]", " ", reddit_data$body)
reddit_data$body <- gsub("[[:digit:]]", " ", reddit_data$body)
reddit_data$body <- gsub("\\s+", " ", reddit_data$body)
# remove stop words
data("stop_words")
reddit_data_clean <- reddit_data %>%
unnest_tokens(word, body) %>%
anti_join(stop_words)
# get sentiment score
reddit_data_sentiment <- reddit_data_clean %>%
inner_join(get_sentiments("nrc")) %>%
count(word, sentiment, sort = TRUE)
# aggregate sentiment for each comment
reddit_sentiment_summary <- reddit_data_sentiment %>%
group_by(sentiment) %>%
summarize(total = sum(n))
# visualize sentiment distribution
ggplot(reddit_sentiment_summary, aes(x = sentiment, y = total, fill = sentiment)) +
geom_bar(stat = "identity") +
theme_minimal() +
labs(title = "Sentiment Distribution in Reddit Comments",
x = "Sentiment", y = "Count")
sentiment_counts <- reddit_sentiment_summary %>% arrange(desc(total)) %>% mutate(percent = (total/sum(total) * 100))
kable(sentiment_counts)
| sentiment | total | percent |
|---|---|---|
| positive | 1268634 | 20.499720 |
| negative | 1014449 | 16.392372 |
| trust | 786426 | 12.707773 |
| fear | 557766 | 9.012881 |
| anticipation | 542107 | 8.759849 |
| anger | 540584 | 8.735239 |
| sadness | 440320 | 7.115083 |
| joy | 413867 | 6.687632 |
| disgust | 381748 | 6.168625 |
| surprise | 242642 | 3.920826 |
set.seed(1)
# Set up parallel processing to speed up random forest training
num_cores <- parallel::detectCores() - 1 # Leave one core free
registerDoParallel(cores = num_cores)
cat("Using", num_cores, "cores for parallel processing\n")
# Create a function to label potential misinformation
# Create a more varied labeling function
label_misinformation <- function(subreddit, text) {
# List of subreddits that might have higher misinformation rates
potential_misinfo_subs <- c("conspiracy", "politics", "worldnews",
"conservative", "liberal")
# Keywords that might indicate misinformation
misinfo_keywords <- c(
"conspiracy", "hoax", "fake", "scam", "they don't want you to know",
"truth revealed", "coverup", "what they won't tell you",
"secret cure", "they're lying", "mainstream media won't report",
"mainstream media lies", "truth they're hiding", "secret agenda", "cover-up",
"they won't tell you", "scientists are hiding", "big pharma doesn't want",
"evidence suppressed", "wake up", "sheeple",
"plandemic", "deep state", "false flag"
)
# Check if text contains any misinformation keywords
contains_keywords <- any(sapply(misinfo_keywords, function(keyword)
grepl(tolower(keyword), tolower(text), fixed = TRUE)))
# Higher probability of misinformation in certain subreddits + keywords
if (subreddit %in% potential_misinfo_subs && contains_keywords) {
return(sample(c(0, 1), 1, prob = c(0.2, 0.8))) # 80% chance of being misinformation
} else if (contains_keywords) {
return(sample(c(0, 1), 1, prob = c(0.5, 0.5))) # 50% chance
} else if (subreddit %in% potential_misinfo_subs) {
return(sample(c(0, 1), 1, prob = c(0.7, 0.3))) # 30% chance
} else {
return(sample(c(0, 1), 1, prob = c(0.9, 0.1))) # 10% chance
}
}
# Limit the dataset size for processing efficiency
sample_size <- min(50000, nrow(reddit_data))
reddit_sample <- reddit_data %>%
sample_n(sample_size)
# Apply preprocessing and labeling
cat("Preprocessing and labeling data...\n")
reddit_processed <- reddit_sample %>%
# Select relevant columns
select(subreddit, body) %>%
rename(text = body) %>% # Rename for consistency
# Apply synthetic labeling
mutate(is_misinformation = mapply(label_misinformation, subreddit, text)) %>%
# Remove empty comments
filter(text != "")
# Print the distribution of our synthetic labels
cat("Synthetic label distribution:\n")
table(reddit_processed$is_misinformation)
# Handle imbalanced classes if necessary
class_counts <- table(reddit_processed$is_misinformation)
min_class_count <- min(class_counts)
# If too imbalanced, sample to create a more balanced dataset
if (min(class_counts) / max(class_counts) < 0.3) {
cat("Balancing dataset due to class imbalance...\n")
# Separate data by class
misinformation <- reddit_processed %>% filter(is_misinformation == 1)
factual <- reddit_processed %>% filter(is_misinformation == 0)
# Determine sample sizes
if (class_counts["0"] > class_counts["1"]) {
# Downsample majority class
factual_sample <- sample_n(factual, min(nrow(factual), nrow(misinformation) * 3))
reddit_balanced <- bind_rows(misinformation, factual_sample)
} else {
# Downsample majority class
misinformation_sample <- sample_n(misinformation, min(nrow(misinformation), nrow(factual) * 3))
reddit_balanced <- bind_rows(misinformation_sample, factual)
}
# Use balanced dataset for further processing
reddit_processed <- reddit_balanced
# Show new distribution
cat("Balanced class distribution:\n")
table(reddit_processed$is_misinformation)
}
# Split data into training and testing sets (80/20 split)
train_indices <- sample(1:nrow(reddit_processed), 0.8 * nrow(reddit_processed))
train_data <- reddit_processed[train_indices, ]
test_data <- reddit_processed[-train_indices, ]
# Feature extraction - TF-IDF approach
cat("Creating vocabulary and TF-IDF features...\n")
# Create vocabulary
it_train <- itoken(train_data$text, progressbar = TRUE)
vocab <- create_vocabulary(it_train)
# Prune vocabulary
vocab <- prune_vocabulary(vocab,
term_count_min = 5,
doc_proportion_max = 0.7,
doc_proportion_min = 0.001)
# Create TF-IDF model
vectorizer <- vocab_vectorizer(vocab)
tfidf_model <- TfIdf$new(norm = "l2", sublinear_tf = TRUE)
# Create document-term matrices with TF-IDF weighting
dtm_train <- create_dtm(it_train, vectorizer)
dtm_tfidf_train <- fit_transform(dtm_train, tfidf_model)
it_test <- itoken(test_data$text, progressbar = TRUE)
dtm_test <- create_dtm(it_test, vectorizer)
dtm_tfidf_test <- transform(dtm_test, tfidf_model)
# Convert to data frames for random forest
X_train <- as.data.frame(as.matrix(dtm_tfidf_train))
y_train <- factor(train_data$is_misinformation)
X_test <- as.data.frame(as.matrix(dtm_tfidf_test))
y_test <- factor(test_data$is_misinformation)
# Feature reduction for Random Forest efficiency
# If there are too many features, we can use PCA or feature selection
feature_count <- ncol(X_train)
cat("Original feature count:", feature_count, "\n")
# If we have too many features, use feature importance from a small forest to select the most important ones
if (feature_count > 500) {
cat("Performing feature selection to reduce dimensionality...\n")
# Train a small random forest for feature selection
small_rf <- randomForest(
x = X_train,
y = y_train,
ntree = 50,
importance = TRUE
)
# Get feature importance
importance_df <- as.data.frame(importance(small_rf))
importance_df$feature <- rownames(importance_df)
# Sort by Mean Decrease in Gini coefficient
importance_df <- importance_df %>%
arrange(desc(MeanDecreaseGini))
# Keep top 500 features
top_features <- head(importance_df$feature, 500)
# Subset the data
X_train <- X_train[, top_features]
X_test <- X_test[, top_features]
cat("Reduced feature count:", ncol(X_train), "\n")
}
# Train Random Forest model
cat("Training Random Forest model...\n")
# Define training control
control <- trainControl(
method = "cv",
number = 5,
verboseIter = TRUE,
allowParallel = TRUE
)
# Define parameter grid for tuning
param_grid <- expand.grid(
mtry = c(floor(sqrt(ncol(X_train))), floor(ncol(X_train)/5), floor(ncol(X_train)/3))
)
# Train model with caret for parameter tuning
rf_model <- train(
x = X_train,
y = y_train,
method = "rf",
ntree = 200, # Fewer trees for speed, increase for better performance
trControl = control,
tuneGrid = param_grid,
importance = TRUE
)
# Print best tuning parameters
cat("Best tuning parameters:\n")
print(rf_model$bestTune)
# Make predictions on test set
cat("Making predictions on test set...\n")
preds <- predict(rf_model, X_test)
# Evaluate model performance
confusion <- confusionMatrix(preds, y_test, positive = "1")
cat("Model evaluation:\n")
print(confusion)
# Get feature importance from the final model
cat("Most important features (top 20):\n")
varImp_df <- varImp(rf_model)$importance
varImp_df$feature <- rownames(varImp_df)
varImp_df <- varImp_df %>%
arrange(desc(Overall)) %>%
head(20)
print(varImp_df)
# Save the model and related objects for future use
saveRDS(rf_model, "reddit_rf_misinformation_model.rds")
saveRDS(vectorizer, "reddit_rf_misinformation_vectorizer.rds")
saveRDS(tfidf_model, "reddit_rf_misinformation_tfidf.rds")
saveRDS(vocab, "reddit_rf_misinformation_vocab.rds")
if (exists("top_features"))
saveRDS(top_features, "reddit_rf_misinformation_top_features.rds")
cat("Random Forest model saved. Load it in the future with:\n")
cat("rf_model <- readRDS('reddit_rf_misinformation_model.rds')\n")
cat("vectorizer <- readRDS('reddit_rf_misinformation_vectorizer.rds')\n")
cat("tfidf_model <- readRDS('reddit_rf_misinformation_tfidf.rds')\n")
cat("vocab <- readRDS('reddit_rf_misinformation_vocab.rds')\n")
if (exists("top_features"))
cat("top_features <- readRDS('reddit_rf_misinformation_top_features.rds')\n")
# Visualization and Presentation of Misinformation Detection Model Results
# Create a directory for visualizations
dir.create("misinformation_visualizations", showWarnings = FALSE)
output_dir <- "misinformation_visualizations/"
# ---- 1. Model Performance Visualizations ----
# Function to create performance visualizations
create_performance_visuals <- function(rf_model, X_test, y_test, output_dir) {
# Get predictions and probabilities
predictions <- predict(rf_model, X_test)
pred_probs <- predict(rf_model, X_test, type = "prob")[,"1"]
# Calculate performance metrics
conf_matrix <- confusionMatrix(predictions, y_test, positive = "1")
# 1.1 Confusion Matrix Heatmap
conf_data <- as.data.frame.matrix(conf_matrix$table)
conf_heatmap <- heatmaply(
conf_data,
dendrogram = "none",
colors = viridis(100),
main = "Confusion Matrix",
xlab = "Predicted",
ylab = "Actual",
margins = c(50, 50),
grid_gap = 0.5,
fontsize_row = 12,
fontsize_col = 12,
cellnote = conf_data,
cellnote_textposition = "middle center",
cellnote_size = 18
)
saveWidget(conf_heatmap, paste0(output_dir, "confusion_matrix.html"))
# 1.2 ROC Curve
pred <- prediction(pred_probs, y_test)
perf <- performance(pred, "tpr", "fpr")
auc <- performance(pred, "auc")@y.values[[1]]
roc_data <- data.frame(
fpr = perf@x.values[[1]],
tpr = perf@y.values[[1]]
)
roc_plot <- ggplot(roc_data, aes(x = fpr, y = tpr)) +
geom_line(color = "#3366CC", size = 1.5) +
geom_abline(slope = 1, intercept = 0, linetype = "dashed", color = "gray50") +
# annotate("text", x = 0.75, y = 0.25,
# label = paste("AUC =", round(auc, 3)),
# size = 5) +
labs(
title = "ROC Curve for Misinformation Detection",
x = "False Positive Rate",
y = "True Positive Rate"
) +
theme_minimal() +
theme(
plot.title = element_text(hjust = 0.5, size = 16),
axis.title = element_text(size = 14),
axis.text = element_text(size = 12)
)
ggsave(paste0(output_dir, "roc_curve.png"), roc_plot, width = 8, height = 6)
# 1.3 Precision-Recall Curve
perf_pr <- performance(pred, "prec", "rec")
# Handle potential NAs
prec_values <- perf_pr@y.values[[1]]
rec_values <- perf_pr@x.values[[1]]
valid_indices <- !is.na(prec_values) & !is.na(rec_values)
pr_data <- data.frame(
recall = rec_values[valid_indices],
precision = prec_values[valid_indices]
)
pr_plot <- ggplot(pr_data, aes(x = recall, y = precision)) +
geom_line(color = "#CC3366", size = 1.5) +
labs(
title = "Precision-Recall Curve",
x = "Recall",
y = "Precision"
) +
theme_minimal() +
theme(
plot.title = element_text(hjust = 0.5, size = 16),
axis.title = element_text(size = 14),
axis.text = element_text(size = 12)
)
ggsave(paste0(output_dir, "precision_recall_curve.png"), pr_plot, width = 8, height = 6)
# 1.4 Model Metrics Summary Table
metrics <- data.frame(
Metric = c("Accuracy", "Sensitivity", "Specificity", "Precision", "F1 Score", "AUC"),
Value = c(
conf_matrix$overall["Accuracy"],
conf_matrix$byClass["Sensitivity"],
conf_matrix$byClass["Specificity"],
conf_matrix$byClass["Precision"],
conf_matrix$byClass["F1"],
auc
)
)
metrics$Value <- round(metrics$Value, 3)
write.csv(metrics, paste0(output_dir, "model_metrics.csv"), row.names = FALSE)
return(list(
conf_matrix = conf_matrix,
auc = auc,
metrics = metrics
))
}
# ---- 2. Feature Importance Visualizations ----
# Function to create feature importance visualizations
create_feature_visuals <- function(rf_model, vocab, output_dir) {
# Get feature importance
imp <- varImp(rf_model)$importance
imp$Feature <- rownames(imp)
# Sort by importance
imp <- imp[order(-imp$Overall), ]
# 2.1 Top 20 Important Features Bar Plot
top_features <- head(imp, 20)
# Reverse the order for better visualization
top_features$Feature <- factor(top_features$Feature, levels = rev(top_features$Feature))
feature_plot <- ggplot(top_features, aes(x = Feature, y = Overall)) +
geom_col(fill = "#5DA5DA") +
coord_flip() +
labs(
title = "Top 20 Features for Misinformation Detection",
x = "Feature (Word/Term)",
y = "Importance Score"
) +
theme_minimal() +
theme(
plot.title = element_text(hjust = 0.5, size = 16),
axis.title = element_text(size = 14),
axis.text = element_text(size = 12)
)
ggsave(paste0(output_dir, "top_features.png"), feature_plot, width = 10, height = 8)
# 2.2 Word Cloud of Important Features
# Prepare data for wordcloud
wordcloud_data <- data.frame(
word = imp$Feature,
freq = imp$Overall
)
# Scale frequencies to be more visually appropriate
wordcloud_data$freq <- wordcloud_data$freq * 100
# Save data for the wordcloud
write.csv(wordcloud_data, paste0(output_dir, "wordcloud_data.csv"), row.names = FALSE)
# Create and save wordcloud (static image version)
wordcloud_plot <- wordcloud2(data = head(wordcloud_data, 100),
size = 0.8,
color = "random-dark",
backgroundColor = "white")
saveWidget(wordcloud_plot, paste0(output_dir, "feature_wordcloud.html"), selfcontained = TRUE)
# 2.3 Feature Correlation Network (for top features)
if(exists("X_train")) {
# Use only top features for the correlation network
top_feature_names <- head(imp$Feature, 30)
# If we have too many observations, sample to speed up correlation calculation
if(nrow(X_train) > 5000) {
set.seed(42)
sample_indices <- sample(1:nrow(X_train), 5000)
X_train_sample <- X_train[sample_indices, top_feature_names]
} else {
X_train_sample <- X_train[, top_feature_names]
}
# Calculate correlation matrix
cor_matrix <- cor(X_train_sample)
# Create correlation plot
png(paste0(output_dir, "feature_correlation.png"), width = 1000, height = 800)
corrplot(cor_matrix, method = "circle", type = "upper",
tl.col = "black", tl.srt = 45, tl.cex = 0.8,
title = "Correlation Between Top Misinformation Features",
mar = c(0, 0, 2, 0))
dev.off()
# Create network graph of correlations (only showing stronger correlations)
# Convert correlation matrix to network
cor_threshold <- 0.3 # Only show stronger correlations
graph_data <- graph_from_adjacency_matrix(
(abs(cor_matrix) > cor_threshold) * cor_matrix,
mode = "undirected",
weighted = TRUE,
diag = FALSE
)
# Plot network
set.seed(42) # For reproducible layout
network_plot <- ggraph(graph_data, layout = "fr") +
geom_edge_link(aes(edge_alpha = abs(weight), edge_width = abs(weight),
color = weight > 0),
show.legend = FALSE) +
scale_edge_color_manual(values = c("FALSE" = "#D55E00", "TRUE" = "#009E73")) +
geom_node_point(size = 5, color = "#0072B2") +
geom_node_text(aes(label = name), repel = TRUE, size = 4) +
labs(title = "Network of Correlated Misinformation Features") +
theme_void() +
theme(plot.title = element_text(hjust = 0.5, size = 16))
ggsave(paste0(output_dir, "feature_network.png"), network_plot, width = 10, height = 8)
}
return(list(
importance = imp
))
}
# ---- 3. Content Analysis Visualizations ----
# Function to analyze comment content and create visualizations
create_content_visuals <- function(reddit_processed, output_dir) {
# 3.1 Distribution of comment lengths
reddit_processed$comment_length <- nchar(reddit_processed$text)
length_plot <- ggplot(reddit_processed, aes(x = comment_length, fill = factor(is_misinformation))) +
geom_density(alpha = 0.6) +
scale_fill_manual(values = c("#3CB371", "#FF6347"),
labels = c("Factual", "Misinformation")) +
labs(
title = "Distribution of Comment Lengths by Class",
x = "Comment Length (Characters)",
y = "Density",
fill = "Classification"
) +
theme_minimal() +
theme(
plot.title = element_text(hjust = 0.5, size = 16),
axis.title = element_text(size = 14),
axis.text = element_text(size = 12),
legend.title = element_text(size = 12),
legend.text = element_text(size = 10)
)
ggsave(paste0(output_dir, "comment_length_distribution.png"), length_plot, width = 10, height = 6)
# 3.2 Misinformation distribution by subreddit
subreddit_counts <- reddit_processed %>%
group_by(subreddit) %>%
summarize(
total_comments = n(),
misinformation_count = sum(is_misinformation),
misinformation_rate = mean(is_misinformation)
) %>%
arrange(desc(misinformation_rate))
# Only include subreddits with enough comments
subreddit_counts <- subreddit_counts %>%
filter(total_comments >= 30)
# Top 20 subreddits by misinformation rate
top_subreddits <- head(subreddit_counts, 20)
# Plot misinformation rate by subreddit
top_subreddits$subreddit <- factor(top_subreddits$subreddit,
levels = top_subreddits$subreddit[order(top_subreddits$misinformation_rate)])
subreddit_plot <- ggplot(top_subreddits, aes(x = subreddit, y = misinformation_rate)) +
geom_col(aes(fill = misinformation_rate)) +
scale_fill_viridis() +
coord_flip() +
labs(
title = "Top Subreddits by Detected Misinformation Rate",
x = "Subreddit",
y = "Misinformation Rate"
) +
theme_minimal() +
theme(
plot.title = element_text(hjust = 0.5, size = 16),
axis.title = element_text(size = 14),
axis.text = element_text(size = 12),
legend.position = "none"
)
ggsave(paste0(output_dir, "subreddit_misinformation_rate.png"), subreddit_plot, width = 10, height = 8)
write.csv(subreddit_counts, paste0(output_dir, "subreddit_misinformation_stats.csv"), row.names = FALSE)
return(list(
subreddit_stats = subreddit_counts
))
}
# ---- 4. Interactive Dashboard for Predictions ----
# Function to create an interactive HTML dashboard for exploring predictions
create_interactive_dashboard <- function(rf_model, vectorizer, tfidf_model,
vocab, test_data, X_test, y_test, output_dir) {
# Get predictions
pred_probs <- predict(rf_model, X_test, type = "prob")[,"1"]
pred_classes <- predict(rf_model, X_test)
# Create results dataset
results_df <- data.frame(
comment = test_data$text,
subreddit = test_data$subreddit,
actual_class = as.character(y_test),
predicted_class = as.character(pred_classes),
confidence = pred_probs,
comment_length = nchar(test_data$text),
correct_prediction = y_test == pred_classes
)
# Create categories for exploration
results_df$prediction_category <- ifelse(
results_df$actual_class == "1" & results_df$predicted_class == "1", "True Positive",
ifelse(results_df$actual_class == "0" & results_df$predicted_class == "0", "True Negative",
ifelse(results_df$actual_class == "1" & results_df$predicted_class == "0", "False Negative",
"False Positive"))
)
# Save full results data for exploration
write.csv(results_df, paste0(output_dir, "prediction_results.csv"), row.names = FALSE)
# Create DataTable for interactive browsing
dt_results <- DT::datatable(
results_df,
options = list(
pageLength = 10,
searchHighlight = TRUE,
order = list(list(4, 'desc')), # Sort by confidence
dom = 'Bfrtip',
buttons = c('csv', 'excel', 'pdf', 'print')
),
filter = 'top',
rownames = FALSE,
caption = htmltools::tags$caption(
style = 'caption-side: top; text-align: center; font-size: 20px;',
'Misinformation Detection Results'
)
) %>%
formatStyle(
'prediction_category',
backgroundColor = styleEqual(
c('True Positive', 'True Negative', 'False Positive', 'False Negative'),
c('#90EE90', '#ADD8E6', '#FFB6C1', '#F0E68C')
)
) %>%
formatStyle(
'confidence',
background = styleColorBar(c(0,1), '#FF9999'),
backgroundSize = '98% 88%',
backgroundRepeat = 'no-repeat',
backgroundPosition = 'center'
)
# Save the interactive table
saveWidget(dt_results, paste0(output_dir, "interactive_results.html"), selfcontained = TRUE)
# Create confidence distribution plot
conf_plot <- ggplot(results_df, aes(x = confidence, fill = prediction_category)) +
geom_density(alpha = 0.7) +
labs(
title = "Distribution of Prediction Confidence by Outcome Category",
x = "Confidence (Probability of Misinformation)",
y = "Density",
fill = "Prediction Outcome"
) +
scale_fill_manual(values = c(
"True Positive" = "#90EE90",
"True Negative" = "#ADD8E6",
"False Positive" = "#FFB6C1",
"False Negative" = "#F0E68C"
)) +
theme_minimal() +
theme(
plot.title = element_text(hjust = 0.5, size = 16),
axis.title = element_text(size = 14),
axis.text = element_text(size = 12),
legend.title = element_text(size = 12),
legend.text = element_text(size = 10)
)
ggsave(paste0(output_dir, "confidence_distribution.png"), conf_plot, width = 10, height = 6)
# High-confidence mistakes analysis
high_conf_mistakes <- results_df %>%
filter(!correct_prediction & confidence > 0.8) %>%
arrange(desc(confidence)) %>%
select(comment, subreddit, actual_class, predicted_class, confidence)
write.csv(high_conf_mistakes, paste0(output_dir, "high_confidence_mistakes.csv"), row.names = FALSE)
return(list(
results = results_df
))
}
# ---- 5. Create Executive Summary Report ----
# Function to generate an executive summary report
create_executive_summary <- function(performance_results, feature_results,
content_results, output_dir) {
# Create a markdown file for the executive summary
sink(paste0(output_dir, "executive_summary.md"))
cat("# Misinformation Detection Model: Executive Summary\n\n")
cat("## Model Performance\n\n")
cat("Our random forest model for detecting misinformation in Reddit comments achieved:\n\n")
cat(paste0("- **Accuracy:** ", round(performance_results$metrics$Value[1], 3), "\n"))
cat(paste0("- **Precision:** ", round(performance_results$metrics$Value[4], 3), "\n"))
cat(paste0("- **Recall:** ", round(performance_results$metrics$Value[2], 3), "\n"))
cat(paste0("- **F1 Score:** ", round(performance_results$metrics$Value[5], 3), "\n"))
cat(paste0("- **AUC:** ", round(performance_results$auc, 3), "\n\n"))
cat("## Key Findings\n\n")
cat("### Top Indicators of Misinformation\n\n")
top_words <- head(feature_results$importance, 10)$Feature
cat("The most predictive terms for identifying misinformation are:\n\n")
for (i in 1:length(top_words)) {
cat(paste0(i, ". **", top_words[i], "**\n"))
}
cat("\n")
if (!is.null(content_results$subreddit_stats)) {
cat("### Subreddits with Highest Misinformation Rates\n\n")
top_subs <- head(content_results$subreddit_stats, 5)
cat("These subreddits had the highest proportion of predicted misinformation:\n\n")
cat("| Subreddit | Misinformation Rate | Total Comments Analyzed |\n")
cat("|-----------|---------------------|-------------------------|\n")
for (i in 1:nrow(top_subs)) {
cat(paste0("| ", top_subs$subreddit[i], " | ",
round(top_subs$misinformation_rate[i], 3), " | ",
top_subs$total_comments[i], " |\n"))
}
cat("\n")
}
cat("## Recommendations\n\n")
cat("Based on our analysis, we recommend:\n\n")
cat("1. **Targeted Moderation:** Focus moderation efforts on subreddits with the highest misinformation rates.\n")
cat("2. **Keyword Monitoring:** Implement automated flagging for content containing top misinformation indicator terms.\n")
cat("3. **Model Refinement:** Consider enriching the model with additional metadata such as user history and post engagement metrics.\n")
cat("4. **Human Review:** Establish a process for human review of high-confidence misinformation predictions.\n\n")
cat("## Limitations\n\n")
cat("Important limitations of this analysis include:\n\n")
cat("- Synthetic labeling was used due to lack of ground truth data, which may introduce biases.\n")
cat("- The model may be more sensitive to certain topics than others.\n")
cat("- Random Forest models can be computationally intensive for very large datasets.\n")
cat("- The context of comments (e.g., sarcasm, quotations) may be misinterpreted.\n\n")
cat("## Next Steps\n\n")
cat("1. Develop a human-in-the-loop system for improving model accuracy over time.\n")
cat("2. Explore deep learning approaches for potentially better performance.\n")
cat("3. Create a browser extension or API to provide real-time misinformation detection.\n")
cat("4. Extend the model to other platforms beyond Reddit.\n")
sink()
return(TRUE)
}
# ---- Main script to run all visualizations ----
# Assuming these objects already exist from previous model training
# rf_model, X_test, y_test, vocab, test_data, etc.
# Replace this section with loading the saved model if needed:
# Load saved model and related objects if not in memory
if (!exists("rf_model")) {
rf_model <- readRDS("reddit_rf_misinformation_model.rds")
vectorizer <- readRDS("reddit_rf_misinformation_vectorizer.rds")
tfidf_model <- readRDS("reddit_rf_misinformation_tfidf.rds")
vocab <- readRDS("reddit_rf_misinformation_vocab.rds")
if (file.exists("reddit_rf_misinformation_top_features.rds")) {
top_features <- readRDS("reddit_rf_misinformation_top_features.rds")
}
}
cat("Creating performance visualizations...\n")
perf_results <- create_performance_visuals(rf_model, X_test, y_test, output_dir)
cat("Creating feature importance visualizations...\n")
feat_results <- create_feature_visuals(rf_model, vocab, output_dir)
cat("Creating content analysis visualizations...\n")
content_results <- create_content_visuals(reddit_processed, output_dir)
cat("Creating interactive dashboard...\n")
dashboard_results <- create_interactive_dashboard(rf_model, vectorizer, tfidf_model,
vocab, test_data, X_test, y_test, output_dir)
cat("Creating executive summary...\n")
create_executive_summary(perf_results, feat_results, content_results, output_dir)
cat("All visualizations and reports have been saved to:", output_dir, "\n")
The random forest model for detecting misinformation in Reddit comments achieved:
| subreddit | total_comments | misinformation_count | misinformation_rate |
|---|---|---|---|
| politics | 49 | 30 | 0.6122449 |
| worldnews | 57 | 33 | 0.5789474 |
| The_Donald | 41 | 20 | 0.4878049 |
| leagueoflegends | 38 | 15 | 0.3947368 |
| pics | 39 | 15 | 0.3846154 |
| trashy | 32 | 10 | 0.3125000 |
| AmItheAsshole | 40 | 11 | 0.2750000 |
| ChapoTrapHouse | 34 | 9 | 0.2647059 |
| news | 42 | 11 | 0.2619048 |
| AskReddit | 31 | 8 | 0.2580645 |
| funny | 30 | 7 | 0.2333333 |
| teenagers | 30 | 6 | 0.2000000 |
| memes | 41 | 8 | 0.1951220 |
| asoiaf | 35 | 6 | 0.1714286 |
| freefolk | 35 | 6 | 0.1714286 |
| dankmemes | 33 | 5 | 0.1515152 |
| apexlegends | 30 | 4 | 0.1333333 |
| hockey | 32 | 4 | 0.1250000 |
| soccer | 33 | 3 | 0.0909091 |
# Create a function to predict misinformation for new Reddit comments
predict_reddit_misinformation <- function(new_comments, new_subreddits = NULL) {
# Preprocess new comments
processed_comments <- tolower(new_comments)
processed_comments <- gsub("[[:punct:]]", " ", new_comments)
processed_comments <- gsub("[[:digit:]]", " ", new_comments)
processed_comments <- gsub("\\s+", " ", new_comments)
# Create document-term matrix with TF-IDF weighting
it_new <- itoken(processed_comments, progressbar = FALSE)
dtm_new <- create_dtm(it_new, vectorizer)
dtm_tfidf_new <- transform(dtm_new, tfidf_model)
# Convert to dataframe
X_new <- as.data.frame(as.matrix(dtm_tfidf_new))
# If we did feature selection earlier, make sure we only use the selected features
if (exists("top_features")) {
missing_cols <- setdiff(top_features, colnames(X_new))
for (col in missing_cols) {
X_new[[col]] <- 0
}
X_new <- X_new[, top_features]
}
# Predict probabilities
pred_probs <- predict(rf_model, X_new, type = "prob")
pred_classes <- predict(rf_model, X_new)
# Create result dataframe
result_df <- data.frame(
comment = new_comments,
probability = pred_probs[,"1"],
prediction = ifelse(pred_classes == 1, "Potential Misinformation", "Likely Factual")
)
# Add subreddit info if provided
if (!is.null(new_subreddits)) {
result_df$subreddit <- new_subreddits
}
return(result_df)
}
# Example usage of the prediction function
example_comments <- c(
"It is a conspiracy theory.",
"This is factual information.",
"The Earth is flat."
)
# Make predictions on example comments
cat("Example predictions:\n")
predictions <- predict_reddit_misinformation(example_comments)
predictions
I am doing some more research, working on my slides and drafting my final paper. I have refined the context of my project to focus solely on Theory of Mind. This makes the project a bit more focused and more directly related to my specific interests in cognitive science. I am finding it fascinating to think about how ToM is employed in a digital space and make hypotheses about how this affects cognition.
For my project, with the help of Claude 3.7 Sonnet, I have created a prototype of my tool. This prototype predicts if a comment contains misinformation and was trained on a Kaggle dataset of Reddit comments. While the prototype is a small approximation of the ideal tool, it is a starting point for the potential of using social media data to learn about social cognition. This tool uses an inductive approach because it will be used to gain insight about social cognition in the digital age using observations gained from the data.
The big question my tool aims to answer is how has social media affected certain social cognitive phenomena?
I hypothesize that, due to the instant nature of social media, people rely more on heuristics rather than going through the entire thinking process they would have if they were processing the information in person. I also hypothesize that Theory of Mind operates over social media but in a slightly different way based on the asynchronous nature of social interactions.
To ensure reproducibility, I used the set.seed function
so the same sequence of random numbers are generated for the test/train
split each time the code is run. Creating a test/train split ensures
that the model is not being trained on the data it will be tested on.
This allows for a more accurate reading on the accuracy of the model.
The tool uses the randomForest to classify Reddit comments
as misinformation or truth. With the test/train split and the random
forest classifier trained on the training data and tested on the test
set, the parameters can be tuned to optimize classifier functioning.
My prototype uses a RandomForest classifier which is a decision tree with bagging for random samples [1]. This approach allows for averaging over the different decision trees to create a more representative statistical model.
Deductive reasoning goes from a theory to evidence to support the theory while inductive reasoning starts with observations and aims to propose a theory [2]. One benefit to the inductive approach is that it allows for discovery of new patterns prevously not classified. However, deductive reasoning tells you exactly what you are looking for whereas inductive reasoning is less specfic and pointed.
References
I used Claude 3.7 Sonnet and asked it: “can you please write me a model in R that predicts if a comment is misinformation? and please adapt it for the Kaggle dataset.”
It responded with a logistic regression model and then recommended randomforests. I then asked it to update the model to use randomforests. It made me a misinformation classification model very quickly and it worked very well. I then asked Claude 3.7 Sonnet to write me a script to visualize the model results. It presented me with a comprehensive R script that even saved png files of the figures.
I am blown away at the capacity of this AI. It is impressive that I can ask it a question in plain English and then it wrote my tool prototype for me.
I need to go through and revise the model and understand the steps, but it is very convenient and helpful to have AI set up a base to work from.
Project Deliverable:
Thinking about cognition and computation and changing oneself
Questions about a tools project:
So far, for my final project, I have selected a dataset, performed EDA and constructed a baseline model. My dataset is a Kaggle dataset of 1 million Reddit comments over 4 subreddits. The variables included in the dataset are subreddit, body (text of the comment), score and controversiality. Score is a calculated value based on the difference between upvotes and downvotes for a comment. Controversiality is a factor that is either 0 (not controversial) or 1 (controversial). A comment is controversial if it has a similar number of upvotes as downvotes. In other words, the score should be close to 0 for a controversial comment. I have done some EDA on the data and learned that 3% of comments in the dataset are considered controversial. I also learned that the score of a comment ranges from -886 to 35619, but is skewed to the right. The most common score a comment has is 1 with 386651 occurrences. Controversial comments have almost a uniform distribution centered around 0. The mean is 0.5027, but the median is 0. It appears as if the minimum value is slightly larger than the maximum value, skewing the data slightly.
For my model, I had to standardize the data. I needed to convert all
of the comments to lowercase and remove punctuation. I also removed stop
words using the stop_words dataset from the
tidytext package because these words do not hold
significant meaning.
I performed sentiment analysis on the dataset using the dataset published in “Crowdsourcing a Word-Emotion Association Lexicon” by Mohammad & Turney (2013). The most common sentiment in the comments is positive, followed by negative then trust.
I have started working on constructing my model. I have identified some words I would expect that a comment with misinformation to contain. Google Gemini and ChatGPT gave me suggestions as well. So far, my list of misinformation key words include: hoax, conspiracy, lie, fake, fraud, outrageous, and insane.
I am currently working on constructing a NLP model to predict misinformation in comments. One roadblock I am currently facing is how to get onto a GPU. Previously, in my data science courses, I had access to a GPU through CWRU OnDemand. I have tried to log on to this, but it appears that I no longer have access to this resource. I plan on asking classmates how they access GPUs for training their models.
While the tool I am constructing will not do exactly what I hope, it is definitely a starting point and prototype for a tool that can take social media data and analyze it to learn about social cognition and misinfomration in a digital world.
Theory of Mind (ToM), or the ability to infer the mental states of others is an integral part of social cognition. It allows humans to understand each other in the absence of having access to another person’s thoughts. Theory of Mind is best understood through an embodied approach. Understanding the mental state of another individual requires incorporating information from having a body in the world and understanding how bodies act (Gallagher, 2020). Thus, creating a computational model that matches human cognition is extremely difficult because computers do not have a mind nor a body. While computers may not be able to execute Theory of Mind in the same way as humans, computational models still serve as an excellent tool for testing hypotheses about the ToM network.
Large language models (LLMs) are incapable of embodied cognition and cannot mirror the cognitive processes that occur in the human brain during Theory of Mind. While LLMs have been shown to perform as well as humans on ToM tasks, the mechanism behind arriving at the correct answer differs, such as in instances of uncertainty (Strachan et al., 2024). Humans have an embodied approach to cognition and can decide how to respond to uncertainty in the world, but LLMs do not have a body in space and are coded to respond in a way that expresses the uncertainty. Rather than making a decision, the computer explains that there is insufficient information to answer the question. It is hypothesized that the disembodiment of LLMs could be one of the reasons they perform so well on false-belief tasks. LLMs do not have their own opinions and beliefs that need to be inhibited, whereas humans need to inhibit their own thoughts and beliefs to answer false-belief questions. Therefore, LLMs, while an excellent tool for studying different aspects of cognition, are an insufficient model for studying Theory of Mind.
However, computational models can serve as an experimental method for testing hypotheses about Theory of Mind on a neurobiological level. The exact brain mechanism for Theory of Mind remains unclear, so models can be created to test how the brain transmits information to understand others’ mental states (Koster-Hale & Saxe, 2013). For instance, one Theory of Mind theory is predictive coding. In this model, information is accumulated based on discrepancies between what is expected and what is occurring. Researchers explain that there is evidence of predictive coding in some of the main ToM brain areas, such as the superior temporal sulcus, the temporoparietal junction and the prefrontal cortex. Computational cognitive scientists can test hypotheses about connectivity between these areas by creating a computational model. Moreover, computational models offer an effective way to repurpose existing data, such as fMRI data from previous ToM studies, to construct and test hypotheses about how this information is transmitted in the brain. Thus, computational modeling is an effective method for studying Theory of Mind neurobiology and brain connectivity.
In sum, while Theory of Mind is a prominent social cognitive phenomenon, the exact neurological mechanism remains unclear. LLMs do not have the capacity to engage in embodied cognition, so they cannot serve as effective models for ToM. However, computational models can be used to test hypotheses about the Theory of Mind brain network. Existing data from previous fMRI studies on ToM can be implemented in training these models so scientists can evaluate hypotheses and examine if any important patterns emerge.
Gallagher, S. (2020). Interaction. In S. Gallagher (Ed.), Action and Interaction. Oxford University Press. https://doi.org/10.1093/oso/9780198846345.003.0006.
Koster-Hale, J., & Saxe, R. (2013). Theory of mind: A neural prediction problem. Neuron, 79(5), 836–848. https://doi.org/10.1016/j.neuron.2013.08.020.
Strachan, J. W. A., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., Gupta, S., Saxena, K., Rufo, A., Panzeri, S., Manzi, G., Graziano, M. S. A., & Becchio, C. (2024). Testing theory of mind in large language models and humans. Nature Human Behaviour, 8(7), 1285–1295. https://doi.org/10.1038/s41562-024-01882-z.
I have decided to pivot (again) for my final project. After reflecting on my personal interests and abilities, I have decided it is not feasible for me to create a computational Theory of Mind model with data that is openly available. Because of this, I was brainstorming different data sources to examine social cognition and thought of social media. There is a plethora of social media data between comments, threads and posts. I wanted to devise a tool that could take this data and extract insights about social cognition.
This is a very big project, so I have decided to take on a single part of it for the purposes of this class. I plan on starting to create this tool by performing sentiment analysis on social media comments to predict misinformation based on vocabulary, like count and (potentially) emotional tone. This is a stepping stone to creation of a tool to analyze social media data to learn about social cognition.
After this is accomplished, the next step would be to obtain data that contains interactions between users and look at responses to comments that are flagged as misinformation. This data would allow scientists to look at belief attribution and examine the beliefs the users must have for the interaction to proceed.
This research is valuable since we have entered a digital age where most communication occurs online. This tool will allow for scientists to examine how interactions differ online compared to in person and how Theory of Mind and belief attribution occurs over the internet.
I found a dataset on Kaggle that contains 1 million Reddit comments across 40 subreddits. I plan on using a NLP model to do sentiment analysis. I will then look at the number of likes these comments received to determine if the comment got a lot of attention or not. I may also try and incorporate emotional tone, but I am not sure yet. Taken together, these predictors should allow me to create a model that predicts if a comment is misinformation, therefore working towards my final goal of tool creation to use social media data to examine social cognition.
The LinkedIn Learning course offered me new insights into R even though I am proficient in the language. For instance, I did not know that R was developed specifically for working with data. I also learned that R is the second most common for use in scholarly articles. The course was a good review on basic R concepts and procedures. I learned about all the different sites that R can be accessed on when I previously thought the only options were R and RStudio. The course also reminded me of all the different types of plots I can create which will be helpful when I am working on data visualizations for my neuroscience capstone project. I think it is really important to have refreshers on skills that one has already learned because we fall into patterns and use heuristics that may not always be the best fit for the task at hand. This course reminded me of different ways to wrangle and display data which I had forgotten about.
I looked at HumanAI, an organization that does machine learning for the arts and humanities. Some of their projects they work on include behavior research with AI and AI generated choreography.
The behavior research project looks at driver behavior to develop interventions to improve transportation safety. The researchers implement gaze analysis to examine distraction in driving. This tool takes data from driving and uses it to help understand distraction for drivers and examines potential interventions for driver safety.
Another project this organization is working on is AI generated choreography. This tool trains AI on dance videos to generate new dance moves. Building on previous projects, this project focuses on duets which are more complex and require more creativity to coordinate two individuals.
These tools can help us look at human cognition by examining behavior and creativity. Human behavior is examined through distracted driving and creativity is examined with duet choreography. These tools take existing data and use it to learn more about human cognition.
Binary Variables
I have made a list of five binary variables, but I could continue the list for forever. There are tons of variables with binary outcomes because anything that is a yes or no response has a binary outcome. I would consider a lot of decision making binary because you either do the thing or you do not. A lot of the world can be seen as binary outcomes, but many argue to avoid thinking this way due to the limiting nature of it. As an individual who tends to see the world in a binary way, sometimes it feels like my entire world works like this when in reality there are intermediates that are sometimes ignored.
A logit variable is the dependent variable in a logistic regression and has a value is between 0 and 1. It estimates the probability of an event as a function of the independent variable.
A probit variable is the dependent variable in a probit model. Probit models are a type of regression that predicts the probability of a binary outcome, usually whether or not a subject has a certain characteristic.
I could create a statistical model of the amount of sleep I obtained and if I will drink coffee or not. I hypothesize that as the amount of sleep I obtain decreases, the likelihood I drink coffee increases. This would be an example of a logistic model. Another logistic model I could make would be on the time I leave for work and if there is traffic. I hypothesize that the chances of traffic increase as I leave closer to rush hour. One example of a probit model would be evaluating whether people are educated or not as a function of income.
I think all of these tools are super cool! It is impressive that cognitive scientists have created these tools to repurpose data we have in our everyday lives. The UCLA Edge Search Engine was easy to use and provided me with more examples of my phrase of interest than I could imagine. The Red Hen clip service worked very well. At first, I did not input my request in the correct format (user error), but once I reformatted the request, I received exactly what I had asked for in a timely manner. The Multidata Pipeline provided me again with more data than I could ever imagine for a 6 second video clip. I am fascinated by how these resources were constructed and amazed at how they can take ordinary media and turn it into valuable data for cognitive scientists. By taking advantage of and repurposing data that already exists, scientists have the ability to spend more time learning about cognition than collecting the data to study. It is so impressive how quickly the tools are able to take an input and produce an output. I feel like we take for granted the work it takes for a tool like these to work so efficiently. I was amazed that within 4 minutes of submitting my request to Multidata, I received a zip file with an enormous amount of data about the video, audio and speech.
Examples of Tools:
I plan on writing my small paper on tools with Theory of Mind. I am interested in how technology can be used to help us learn about this phenomenon. This paper will also aid in background knowledge for my final project which looks at Theory of Mind as a computational model.
I have been struggling with how to proceed with my final project because the scope was too large. I have refined my final project to make it more manageable. I will be creating a Theory of Mind computational model with executive function aspects. This differs from my original idea by omitting the brain aspect and requirement for fMRI data.
I have found an open source Theory of Mind false-belief task dataset from 2023. This dataset contains results of participant accuracy along with other information.
Theory of Mind Booklet Task & Open Dataset
I need to clean the data and subset it for the data I want for my project. This dataset has results from 3 different projects and I plan on only using results from 1.
Deliverable
I will create computational Theory of Mind modelu sing reinforcement learning that can be manipulated to examine the role of executive function in ToM.
Plan
Reference
Sotomayor-Enriquez, K., Gweon, H., Saxe, R., & Richardson, H. (2023, August 28). Open dataset of theory of mind reasoning in early to middle childhood. Retrieved from psyarxiv.com/gczp9
Deliverable
I will create computational Theory of Mind model based on the mentalizing brain network and reinforcement learning that can be manipulated to learn about ToM and executive dysfunction.
Plan
I think the best way to create this model is in Python, since I will be attempting to construct a deep neural network to mimic brain area interactions.
I asked ChatGPT to help me devise a model architecture since I have never created a project like this before. Here are the summarized results:
I also asked ChatGPT to write pseudocode for this model as a starting point for me.
I will use publicly available data to train my model. There is a dataset on NeuroVault from a fMRI study on Theory of Mind that I can use to biologically ground my model. There is another dataset from OpenNeuro on a ToM task with responses that I can use to train my model.
Steps
Kaggle Student Mental Health Dataset
To begin my data analysis, I loaded in the necessary libraries for
EDA and data visualization, such ash tidyverse and
ggplot2. I then cleaned up the dataframe to make it more
user friendly by renaming variables and converting character variables
into factors. I printed a preview of the dataset using
str() and printed summary statistics.
Next, I did some exploratory data analysis and looked at the frequency of different entries in the data such as the number of females versus males, the number of individuals that reported having anxiety, the number of individuals that reported having depression and the frequency of different GPA ranges.
Then, I created contingency tables for relationships between variables I was interested in. These included the relationship between GPA and anxiety and GPA and depression. I ran chi-square tests on both of these associations and found that there appears to be a significant relationship between depression and GPA (p = 0.061).
Some other correlations I would be interested in exploring are the relationship between age and mental health condition, year in school and mental health condition and gender and mental health condition.
# load in packages
library(tidyverse)
library(dplyr)
library(ggplot2)
library(ggpubr)
library(kableExtra)
# read in data
mental_health_df <- read.csv("Student Mental health.csv")
# tidy data
mental_health_df <- mental_health_df %>%
rename(gender = Choose.your.gender,
age = Age,
major = What.is.your.course.,
year = Your.current.year.of.Study,
gpa = What.is.your.CGPA.,
marital_status = Marital.status,
depression = Do.you.have.Depression.,
anxiety = Do.you.have.Anxiety.,
panic_attack = Do.you.have.Panic.attack.,
treatment = Did.you.seek.any.specialist.for.a.treatment.)
mental_health_df$year <- str_to_lower(mental_health_df$year)
mental_health_df$gpa <- str_trim(mental_health_df$gpa)
mental_health_df <- mental_health_df %>%
mutate(gender = as.factor(gender),
year = as.factor(year),
gpa = as.factor(gpa),
marital_status = as.factor(marital_status),
depression = as.factor(depression),
anxiety = as.factor(anxiety),
panic_attack = as.factor(panic_attack),
treatment = as.factor(treatment))
str(mental_health_df)
# summary statistics
summary(mental_health_df)
# EDA
age_plot <- ggplot(mental_health_df) +
geom_bar(aes(x = age), fill = "cornflowerblue")+
guides(x = guide_axis(angle = 45)) +
labs(title = "Age")
gender_plot <- ggplot(mental_health_df) +
geom_bar(aes(x = gender), fill = "seagreen")+
guides(x = guide_axis(angle = 45)) +
labs(title = "Gender")
year_plot <- ggplot(mental_health_df) +
geom_bar(aes(x = year), fill = "purple4")+
guides(x = guide_axis(angle = 45)) +
labs(title = "Year in School")
gpa_plot <- ggplot(mental_health_df) +
geom_bar(aes(x = gpa), fill = "pink2") +
guides(x = guide_axis(angle = 45)) +
labs(title = "Cumulative GPA")
anxiety_plot <- ggplot(mental_health_df) +
geom_bar(aes(x = anxiety), fill = "orchid")+
guides(x = guide_axis(angle = 45)) +
labs(title = "Anxiety")
depression_plot <- ggplot(mental_health_df) +
geom_bar(aes(x = depression), fill = "aquamarine3")+
guides(x = guide_axis(angle = 45)) +
labs(title = "Depression")
mental_health_eda_plot <- ggarrange(age_plot, year_plot, gpa_plot,gender_plot,
anxiety_plot, depression_plot)
annotate_figure(mental_health_eda_plot, top = text_grob("Mental Health EDA",
color = "black", face = "bold", size = 14))
# chi square analysis
anxiety_table <- table(mental_health_df$gpa, mental_health_df$anxiety)
kable(anxiety_table, caption = "GPA and Anxiety Contingency Table")
chisq.test(mental_health_df$gpa, mental_health_df$anxiety)
depression_table <- table(mental_health_df$gpa, mental_health_df$depression)
kable(depression_table, caption = "GPA and Depression Contingency Table")
chisq.test(mental_health_df$gpa, mental_health_df$depression)
## 'data.frame': 101 obs. of 11 variables:
## $ Timestamp : chr "8/7/2020 12:02" "8/7/2020 12:04" "8/7/2020 12:05" "8/7/2020 12:06" ...
## $ gender : Factor w/ 2 levels "Female","Male": 1 2 2 1 2 2 1 1 1 2 ...
## $ age : int 18 21 19 22 23 19 23 18 19 18 ...
## $ major : chr "Engineering" "Islamic education" "BIT" "Laws" ...
## $ year : Factor w/ 4 levels "year 1","year 2",..: 1 2 1 3 4 2 2 1 2 1 ...
## $ gpa : Factor w/ 5 levels "0 - 1.99","2.00 - 2.49",..: 4 4 4 4 4 5 5 5 3 5 ...
## $ marital_status: Factor w/ 2 levels "No","Yes": 1 1 1 2 1 1 2 1 1 1 ...
## $ depression : Factor w/ 2 levels "No","Yes": 2 1 2 2 1 1 2 1 1 1 ...
## $ anxiety : Factor w/ 2 levels "No","Yes": 1 2 2 1 1 1 1 2 1 2 ...
## $ panic_attack : Factor w/ 2 levels "No","Yes": 2 1 2 1 1 2 2 1 1 2 ...
## $ treatment : Factor w/ 2 levels "No","Yes": 1 1 1 1 1 1 1 1 1 1 ...
## Timestamp gender age major year
## Length:101 Female:75 Min. :18.00 Length:101 year 1:43
## Class :character Male :26 1st Qu.:18.00 Class :character year 2:26
## Mode :character Median :19.00 Mode :character year 3:24
## Mean :20.53 year 4: 8
## 3rd Qu.:23.00
## Max. :24.00
## NA's :1
## gpa marital_status depression anxiety panic_attack treatment
## 0 - 1.99 : 4 No :85 No :66 No :67 No :68 No :95
## 2.00 - 2.49: 2 Yes:16 Yes:35 Yes:34 Yes:33 Yes: 6
## 2.50 - 2.99: 4
## 3.00 - 3.49:43
## 3.50 - 4.00:48
##
##
| No | Yes | |
|---|---|---|
| 0 - 1.99 | 4 | 0 |
| 2.00 - 2.49 | 2 | 0 |
| 2.50 - 2.99 | 3 | 1 |
| 3.00 - 3.49 | 28 | 15 |
| 3.50 - 4.00 | 30 | 18 |
##
## Pearson's Chi-squared test
##
## data: mental_health_df$gpa and mental_health_df$anxiety
## X-squared = 3.5243, df = 4, p-value = 0.4742
| No | Yes | |
|---|---|---|
| 0 - 1.99 | 4 | 0 |
| 2.00 - 2.49 | 2 | 0 |
| 2.50 - 2.99 | 1 | 3 |
| 3.00 - 3.49 | 24 | 19 |
| 3.50 - 4.00 | 35 | 13 |
##
## Pearson's Chi-squared test
##
## data: mental_health_df$gpa and mental_health_df$depression
## X-squared = 8.9975, df = 4, p-value = 0.06116
I selected a model project that combines Theory of Mind and executive dysfunction because it combines my interests in developmental neurobiology and social cognition along with my curiousity about modeling cognitive phenomena. I would like to create a model that can be manipulated to explore executive dysfunction to examine how this affects performance on false-belief tasks. I am aware that computational models perform well on false-belief tasks, but individuals with impaired executive function do not. I am interested in exploring specifically why this difference arises between neurotypical individuals and individuals with neurodevelopmental disorders.
I plan on creating a computational model of Theory of Mind that can be modified to examine executive dysfunction. Based on the Deep Research output, I would like to create a hybrid computational model that uses neurocognitive stuff. It will also have parameters for executive functions like inhibitory control, in which I am very interested. The modular design of a brain-inspired hybrid model allows for manipulations, such as with executive functions.
This model should have components that mirror brain areas involved in ToM such as perception (STS), belief inference (TPJ) and decision making/executive control (mPFC). The model should also have mechanisms for inhibitory control that can be manipulated.
I am not very well-versed in creating models, so I need to do some more research on how to execute this. I am confident on my ToM knowledge and my understanding of ToM with executive dysfunction. I just need to translate this knowledge into a computational setting. My plan to do this is to work with AI to help me understand how to build a computational model and then take this baseline understanding to adapt a model for my research needs.
The phrase I decided to explore is “out of my system.” I hypothesize that a speaker will make a gesture outward from him/herself while saying this phrase.
Five examples of “out of my system” include:
I find it interesting that in my five examples, there are slight differences in the way the gesture is carried out. Each example has the speaker moving their hands away from themself but in slightly different ways. This makes me wonder if the context the phrase is being spoken in affects the way the gesture is carried out.
This work can be attributed to my name.
Trust_P1_YOU
as the dependent variable and Trust_Dictator_P1_YOU as the
independent variable##
## Call:
## lm(formula = Trust_P1_YOU ~ Trust_Dictator_P1_YOU, data = bathumTrustClean)
##
## Residuals:
## Min 1Q Median 3Q Max
## -2.2676 -1.3983 -0.3983 0.6017 3.6017
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 1.39832 0.12322 11.348 <2e-16 ***
## Trust_Dictator_P1_YOU 0.08692 0.06757 1.287 0.2
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 1.694 on 228 degrees of freedom
## Multiple R-squared: 0.007207, Adjusted R-squared: 0.002853
## F-statistic: 1.655 on 1 and 228 DF, p-value: 0.1996
The Deep Research output from my question was very thorough. I am unsure why I asked about the theory and neural underpinnings of ToM in my query, as I am very familiar with these aspects. However, the report sections on these topics did match with my prior knowledge, which is a good thing. I was mainly interested in learning about computational ToM models and how I could implement one to examine executive dysfunction. The results offered five different computational strategies with its respective strengths and weaknesses for modeling ToM. Then, based on the different approaches, the output offered me a suggestion on the optimal way to create a computational ToM model. Its suggestion was to create a hybrid ToM model using a neural network with a modular structure. This was the idea I had when I came up with my project idea, but I had more of a theoretical sense of what I wanted to do. The Deep Research output provided me with context for the idea.
The executive summary at the beginning of the document does a nice job outlining the research and gives a brief synopsis of the recommendations based on my question. The recommendations section is a thorough conceptual outline of how Deep Research suggests I create my model. Moreover, the research offered me sources to look at more closely to help me with my project.
Overall, I found the Deep Research output helpful and will use its insights as a starting point for my final project resarch.
Prompt that was run in Deep Research
“Act as a cognitive scientist who is a specialist in social cognition, especially in Theory of Mind. Theory of Mind is a social cognitive phenomenon that explains the ability to understand the mental states of others. I am seeking to create a Theory of Mind computational model that simulates how the brain functions during this phenomenon. Generate a comprehensive, detailed research report using clear, direct academic English about Theory of Mind computational models, brain areas activated in Theory of Mind and how these areas differ in instances of executive dysfunction. Analyze different computational theories that explain Theory of Mind and create a table comparing them. Please suggest the optimal theory for creating a Theory of Mind model and make recommendations for alternative models as well. Give recommendations on how to construct an effective Theory of Mind model that can be manipulated to explore executive dysfunction and Theory of Mind. Examine the optimal way to construct a computational Theory of Mind model that can be manipulated to study Theory of Mind and executive dysfunction. The report should be separated with headers, include an executive summary, recommendations and include detailed citations. Please concentrate on peer-reviewed articles, but include excellent research from other sources as you see fit.”
Questions Deep Research asked:
“Thank you for your request. Before I begin, could you specify the preferred citation style for the report (e.g., APA, MLA, Chicago)? Also, do you have any specific computational modeling approaches in mind (e.g., Bayesian models, neural networks, symbolic AI), or should I explore a broad range of approaches? Lastly, do you need recommendations on software or programming tools for building the model?”
Answers to these questions:
“Use APA citation style. Explore a broad range of modeling approaches and compare their merits and demerits for this project. Please include recommendations as suitable.”
PiCT-FETE
P: personal (who am I and who are you/the AI)
C: context (what’s the background)
T: what’s the task (be elaborate and specific)
F: format (give example if possible)
T: tone (how do you want it written (e.g. clear, direct, academic English)
“Act as a cognitive scientist who is a specialist in social cognition, especially in Theory of Mind. I am seeking to create a Theory of Mind computational model that simulates how the brain functions during this phenomenon. Generate a comprehensive, detailed research report about Theory of Mind computational models, brain areas activated in Theory of Mind and how these areas differ in instances of executive dysfunction. Analyze different computational theories that explain Theory of Mind and select the best theory. Give recommendations on how to construct an effective Theory of Mind model that can be manipulated to explore executive dysfunction and Theory of Mind. Make recommendations for alternative models as well. The report will examine the optimal way to construct a computational Theory of Mind model that can be manipulated to study Theory of Mind and executive dysfunction. The report should be separated with headers and include a full bibliography. Please use peer-reviewed journal articles in your research.”
Answers to potential follow-up questions:
I am still exploring CQPweb, so I made my query fairly simple for this assignment. I understand how to create a query using CQPweb syntax, but I am not very comfortable with devising linguistic patterns to explore. I do not have a strong background in lingustics, so this part is a challenge for me. My query for this week used a simple grammatical structure of noun + verb + adverb. One example of this structure is “I ran quickly.”
The CQP syntax to query for conditional statements is: [pos = “NN”] [pos = “VB”] [pos = “RB”]
The results from this query include:
One example from the query results is:
“In some respects her story is like that of another doctor who in a moment of thoughtless fury saw her career go up in flames” (t__f648d2e4_1669_11e7_9302_089e01ba0770).
“Career go up” is an example of a phrase that consists of a noun (career), verb (go) and adverb (up). While not a complex linguistic pattern, noun + verb + adverb is a common grammatical structure used in everyday life.
One note on the CQPweb query I used is that I used the part of speech tags for noun singular or mass (NN), verb in the base form (VB), and adverb (RB). If I wanted to create a more specific query, I could use different noun and verb forms, such as plural (NNS) or proper nouns (NNP(S)) or different forms of verbs (past tense (VBD), present/past participle (VBG/VBN). I could have also created a more specific query by making different components optional (?) or offering alternatives (…|…).
Theory of Mind (ToM) is the ability to infer the mental states of others. The term was coined by Premack and Woodruff in the seminal paper, Does the chimpanzee have a theory of mind? The researchers explain that humans can infer purpose, knowledge, belief and pretend (Premack & Woodruff, 1978). These abilities were assessed in a chimpanzee through a handful of tasks. While this study is not ecologically relevant, the researchers provide a definition and explanation of Theory of Mind that is still referenced today. To assess ToM, false-belief tests are typically administered. One example of these tests is the Sally-Anne task where participants are asked to infer the mental state of one of the characters. Children tend to pass the assessment around age 4. Since ToM is an important social cognitive phenomenon, there are a variety of theories describing the mechanism behind it. These include Theory Theory (TT), Simulation Theory (ST), Interaction Theory (IT), the Theory of Mind Mechanism (ToMM).
Different cognitive Theory of Mind theories provide unique lenses to examine ToM. Theory Theory explains that people infer the mental states of others by implementing folk psychology. In other words, we read others’ minds by using a set of causal laws that relate to inner states, similar to scientific theory (Gallese & Goldman, 1998). TT arises because other peoples’ mental states are unobservable. In contrast to TT, Simulation Theory posits that people create a mental simulation of a situation to infer the mental states of others. To do this, an individual creates a mental simulation in their own mind to infer what another person is experiencing which requires having previous personal experience of the situation. ST is supported by the discovery of mirror neurons. Mirror neurons were first discovered in the monkey premotor cortex and fire when an action is performed as well as when it is observed. Unlike TT and ST, Interaction Theory utilizes 4E cognition to explain ToM. IT explains that humans have an embodied and embedded experience of the world that allows us to understand, rather than infer, the mental states of others (Gallagher, 2020). In other words, humans use the context of the world around them to understand the mental states of others. Thus, TT, ST and IT each offer distinct perspectives to understand ToM.
A neurobiological approach to Theory of Mind is the Theory of Mind Mechanism. This theory explains that humans are born with mental machinery to infer others’ mental states, so ToM is an innate ability (Leslie et al., 2004). ToMM posits that one’s hypothesis about another person’s belief is the same as one’s own, which is referred to as the true belief default. The true belief default needs to be inhibited to pass a false-belief test because the true belief is not what the deceived character will believe. Therefore, ToMM offers a neurobiological lens to explore ToM.
In conclusion, Theory of Mind is a complex social cognitive phenomenon that has been theorized about by many. Different approaches to ToM provide different frameworks for understanding how it develops and when it is not functioning properly. These lenses also offer unique pathways for tools and models to study ToM.
Gallagher, S. (2020). Interaction. In S. Gallagher (Ed.), Action and Interaction. Oxford University Press. https://doi.org/10.1093/oso/9780198846345.003.0006.
Gallese, V., & Goldman, A. (1998). Mirror neurons and the simulation theory of mind-reading. Trends in Cognitive Sciences, 2(12). https://doi.org/10.1016/S1364-6613(98)01262-5.
Leslie, A. M., Friedman, O., & German, T. P. (2004). Core mechanisms in ‘theory of mind.’ Trends in Cognitive Sciences, 8(12), 528–533. https://doi.org/10.1016/j.tics.2004.10.001.
Premack, D., & Woodruff, G. (1978). Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4), 515–526.https://doi.org/10.1017/S0140525X00076512.
I feel very confident in my ability of creating visualizations of
different models from my background completing DSCI 351/352/353. I am
experienced in ggplot2 and other visualization libraries
for other statistical models, such as ggraph for graph
models.
In DSCI 352, I created a variety of visualizations of my data during
the exploratory data analysis phase using scatterplots, bar graphs, box
plots and correlation matrices. This project also required me to create
a graphical representation of the data which I did using the
igraph package in R. Then, I attempted to create a graph
neural network to predict connectivity between nodes.
When it comes to creating statistical models, I am familiar with linear and logistic regressions, LDA and QDA, and k nearest neighbors. I am also comfortable with different types of correlation analysis and the implementation of different correlation coefficients. Additionally, I have had an introduction to the Bayes classifier, SVM and neural networks, but am not very comfortable with these models.
One concept I would like to learn more about is the workings of neural networks. I was given an elementary introduction to neural networks in DSCI 353, but do not understand the math behind them. One step to do this is refresh my linear algebra knowledge and find videos online that explain the math of neural networks.
Another aspect of models I would like to work on is advancing my knowledge on graphs. In fact, I have found some papers written by my grandfather, Dr. Bob Wilkov, wrote in the 1970s about graph networks and how they function which I would like to study. While his papers are in the context of electrical engineering, the concept of how graph networks function and the math behind them.
I have decided I want to pivot for my final project. Initially, I wanted to create a tool for studying Theory of Mind (ToM), but have decided to shift to create a model of Theory of Mind in ASD or ADHD. This new project idea continues my studies of Theory of Mind, but I feel like I am more capable of making this model than creating a tool to study ToM.
To create my model, I plan on researching models of ToM and ToM models of other disorders. I then will examine the neural correlates of ToM deficits in ASD or ADHD which I will use to construct a model of ToM that takes these deficits into account.
1. A research question or hypothesis having to do with cognition and computation.
How does the brain communicate to produce Theory of Mind abilities? How can the Theory of Mind network be modeled? How does this network deviate in individuals with neurodevelopmental disorders?
2. Data that could help test this question or hypothesis. Does it exist? Where would you get it?
fMRI data from previous studies can be used to determine the brain areas implicated. The data would need to be publicly available. These studies offer insight into which brain areas are recruited during Theory of Mind tasks. The data would provide a framework for creating a model of Theory of Mind.
3. Thoughts about how to tag (if tagging is needed) and search this data to produce a comma-separated-value file of results.
fMRI data is already analyzed and interpreted by the researchers from the initial studies, so I do not think tagging would be needed. A meta-analysis of previously published results could be conducted to determine the brain areas most activated during Theory of Mind tasks which would provide us with a CSV file of results.
4. Calculations that you would do in R on the data that is contained in this comma-separated-value file so as to produce some results.
The data in the CSV file can be run through statistical tests such as Heges g to see the size effect between control and experimental groups. The \(I^2\) statistic can be used to assess the heterogeneity of the conditions.
5. A visualization of these results.
The results can be visualized in a forest plot to examine heterogeneity between experiments to determine the strength of the meta-analysis and the results. After the statistical analysis, we can create a graph representing proposed connectivity between brain areas activated during ToM. After completing this, the same procedure can be repeated but for experiments with individuals with neurodevelopmental disorders. The results from that meta-analysis can be compared with these results and then the graph can be altered to match our hypothesis of the ToM pathway in individuals with neurodevelopmental disorders.
One interesting linguistic pattern is a conditional statement. Conditional statements are in the format, “if X is true, then Y is true” which can also be explained as “if [condition], then [result].” This linguistic pattern demonstrates a relationship between two ideas.
The CQP syntax to query for conditional statements is: [lemma = “if”] []* [pos = “V.*”]
The results from this query include:
One example from the query results is:
“Under Oklahoma Law, if someone dies during the Commission of a felony, all suspects could be charged with murder even if they didn’t actually kill someone” (t__f648d2e4_1669_11e7_9302_089e01ba0770).
In this sentence, the conditional statement begins with “if someone dies” and then is followed by “all suspects could be charged with murder.” Thus, a death is the condition and the suspects being charged with murder is the result.
“Please research the history, theories and neural correlates of Theory of Mind. I am interested in learning about the emergence of the theory and different interpretations of it. Moreover, I would like to know the brain areas that are activated during Theory of Mind tasks. Please use peer-reviewed articles in your research and produce a report.”
Answers to potential follow-up questions:
I followed the steps in the syllabus to run a chi-square test for the
two way table of bat$Trust_P1_YOU and
bat$Donation_P1_YOU. The result returns a p-value of
1.746e-15 which is a very small number. This means that the correlation
between the two variables is statistically significant and there is a
very small chance that the null hypothesis is correct, so we can reject
it. In other words, there is a very small chance that the correlation
observed is due to chance alone and that there is a correlation between
bat$Trust_P1_YOU and bat$Donation_P1_YOU.
The p-value represents the smallest level of significance needed to reject the null hypothesis. It is the probability of incorrectly rejecting the null hypothesis, so if this value is small, we can be confident that we can reject the null hypothesis.
A chi-square test is a type of hypothesis test that examines the difference between observed and expected values. It is commonly used with contingency tables. A large chi-square value means that there are large deviations between the observed and expected values, the degree of freedom is the number of observations - 1 and then the p-value tells if the difference between observed and expected is large enough to attribute the difference to a correlation and not due to chance.
\[\chi^2 = \sum \frac {(O_i - E_i)^2}{E_i}\]
In this example, we are testing to see if the donation values
(observed) are significantly different from the trust values (expected).
Based on the chi-square result, we can reject the null hypothesis that
the relationship between bat$Trust_P1_YOU and
bat$Donation_P1_YOU is due to chance and conclude that
there is a significant relationship between these variables.
Reflecting on my statistical knowledge, I think I need to refresh myself on statistical principles. I believe my use of statistical tests has become automatic and I do not think critically before performing statistical analysis.
# two way table for chi square test
send_money <- table(bat$Trust_P1_YOU, bat$Donation_P1_YOU)
# chi square
chisq.test(send_money)
##
## Pearson's Chi-squared test
##
## data: send_money
## X-squared = 126.22, df = 25, p-value = 1.746e-15
To be completely honest, I am still struggling with the idea of material anchors. I read the article by Edwin Hutchins on material anchors and asked NotebookLM to explain it as well. From my understanding, material anchors are items in the physical environment that are incorporated into mental concepts. These physical objects stabilize conceptual ideas.
Material anchors reduce cognitive load for mental computation. This is because part of the concept (the material anchor) is a physical entity and does not need to be constructed in one’s mind. Material anchors also help with complex reasoning because of the cognitive offloading to a physical entity.
One example of this could be with the clock example. Clocks are a material anchor for time by blending a physical object (clock) to an abstract concept (time). Clocks have hands that move which are related to the passage of time. Clocks reduce cognitive load by making time easier. It is easier to think about tasks in minutes or hours, or the passage of minutes or hours, when related to hands moving on a clock. Since mental capacity is not being used to keep track of time, the mind is free to do other tasks.
Material anchors relate to 4E cognition because implementing these anchors requires an embedded and enactive approach. We are interacting with our environment and using this experience to create conceptual blends that include material anchors. This process requires observing something in the environment and then incorporating it into a mental representation to create a conceptual blend.
I feel like material anchors in conceptual blends are implemented in metaphor. When we say things like “life is a highway” we are using a highway as a material anchor for life going in a direction and time moving quickly. This blends highway with life to create a conceptual blend of time moving quickly during one’s life.
One question I have is if context-dependent memory is a form of a conceptual blend. For instance, if I am in the kitchen and realize I need a recipe from my room, I go to my room and forget what I went in there to retrieve. I return to the kitchen and remember that I needed the recipe from my room to continue cooking dinner. Is the kitchen a material anchor for remembering I need a recipe? To me, this appears like an application of the Method of Loci, but I am unsure. This example lacks sequential mapping but employs a physical space to map a concept (needing a recipe). I am blending a memory to do a task with the physical space to which it is related.
Since I plan to do my semester project on tools and Theory of Mind, I feel it is appropriate to write my theory paper on Theory of Mind as well. I have learned about this topic in my cognitive science classes but I think it would be beneficial if I read some of the important papers such as the 1978 paper by Premack and Woodruff defining Theory of Mind. I am also interested in Theory of Mind computational theory but am unsure if this falls into the theory or models category. I would like to learn more about how Theory of Mind is represented in computational models and why they are able to outperform humans on false belief tasks. Based on my interests, I plan on writing a theory paper about Theory of Mind and potentially Theory of Mind computational theory and then my models paper on Theory of Mind models. These papers will provide me with sufficient background information to embark on my project to hopefully create a Theory of Mind dataset where humans outperform artificial intelligence.
Trust_P1_YOUggplot(data = bat) +
geom_bar(mapping = aes(x = Trust_P1_YOU), fill = "lightpink") +
labs(title = "Bar Chart of Trust_P1_YOU")
Donation_P1_YOUmeanDonation_P1_YOU = mean(bat$Donation_P1_YOU, na.rm = TRUE)
medianDonation_P1_YOU = median(bat$Donation_P1_YOU, na.rm = TRUE)
varianceDonation_P1_YOU = var(bat$Donation_P1_YOU, na.rm = TRUE)
sdDonation_P1_YOU = sqrt(var(bat$Donation_P1_YOU, na.rm = TRUE))
cat("The mean of Donation_P1_YOU is", meanDonation_P1_YOU,
"\nThe median of Donation_P1_YOU is", medianDonation_P1_YOU,
"\nThe variance of Donation_P1_YOU is", varianceDonation_P1_YOU,
"\nThe standard deviation of Donation_P1_YOU is", sdDonation_P1_YOU, "\n")
## The mean of Donation_P1_YOU is 1.375
## The median of Donation_P1_YOU is 1
## The variance of Donation_P1_YOU is 2.966398
## The standard deviation of Donation_P1_YOU is 1.722323
na.rm = TRUE in the R code.na.rm is a logical value that asks whether NA values
should be excluded during calculations. When we say
na.rm = TRUE, this means that we want R to exclude NA
values when calculating the metric. Measures of central tendency cannot
be calculated with NA values, so they need to be excluded before doing
the calculation.
# create data
junk1 <- read_csv("ID, length, width
1, 7, 49
2, 2, 4
3, 12, 144
4, 10, 100
5, 1, 1
6, 6, 36
")
# plot
ggplot(junk1) +
geom_point(aes(x = length, y = width), color = "cornflowerblue") +
labs(title = "Graph of Width vs Length")
I completed the Generative AI Education course and submitted the Student Survey. I think the course did a great job setting up the basics of Gen AI. As a data science minor, my previous courses have covered how large language models function and the importance of data privacy and security. Because of this, the course was a surface-level overview of Gen AI. Even though I had previous knowledge of artificial intelligence, the course taught me how to use Gen AI for academic purposes. I did not know how to formulate an effective prompt before this course. Overall, I think the course did an excellent job providing students with an overview of Gen AI and how to use it effectively and appropriately for academic purposes.
The Edge Search Engine is a fascinating tool. I have never seen something like it before. It is an appealing tool to construct a dataset to explore questions about cognition. For instance, we could ask and answer questions about the frequency of phrases used or metaphors in the media. It could also be helpful for looking at gestures associated with language and understanding frames and reframing.
If I am being completely honest, I am struggling to come up with topics for my effort on tools. Tools are supposed to help us learn about human cognition from extant data. To do this, there has to be a phenomenon of interest identified, a dataset acquired or created and then the tool to be constructed. The specific aspects of higher order cognition I am interested in include social cognition and executive functioning. Tools can be used to help us understand cognition and analyze data. If there was a plethora of data on executive function tasks from typically developing individuals and individuals with deficits in executive function, a tool could be created to analyze the differences between the two groups and see what factors play a role in these differences.
Another way computational tools could be used to explore cognition is using video data of social interactions and using statistical analysis to understand identifying features of social interactions. For example, looking at eye gaze or body language when somebody is speaking. This would help us identify key components of social interactions and maybe help us identify which aspect humans are most sensitive to. Like, if body language and tone are inconsistent, then is the person able to interpret the correct meaning? This would be a good question for a tool to help us with since there is a lot of data (video footage) of people having social interactions that a computer can help us statistically analyze for results.
I would be interested in exploring Theory of Mind and computation because it combines my interests in social cognition and data science. I would like to incorporate aspects seen in neurodevelopmental disorders to further my understanding of how ToM devaites in this population.
I have a special interest in autism spectrum disorder (ASD) and other neurodevelopmental disorders that impact higher cognition because my sister and cousins have ADHD and I also have a cousin with ASD. Because of my interests, I work in Dr. Heather Broihier’s lab which studies genes related to autism and synaptic refinement, volunteer at Monarch School for Autism, and am enrolled in a cognitive science independent study to explore the effects of ASD on cognition in depth. Because of my interests, I am intrigued by the prospect of computational tools that allow for analysis of executive function or social cognition, two of the major cognitive science topics that interest me. I think it would be fascinating if I could take already published data, manipulate it and come to conclusions about a facet of executive function or social cognition from this.
Last semester, I looked at Theory of Mind (ToM) and its current assessments. I found that individuals with ASD and ADHD both demonstrate deficits in this ability but the etiology differs. I also read an article that examined Large Language Models’ (LLM) abilities to complete ToM tasks. The article found that LLMs could outperform human subjects in some instances. Not only did this article further my interest in the intersection of data science and cognition but it had me questioning the role LLMs should play in cognitive science research. After the first class, I am interested in the idea of using data science as a tool to explore topics in cognitive science, but am hesitant about computational models of highly embodied and embedded processes.
I am majoring in neuroscience and cognitive science with a minor in applied data science. I have taken all of the courses for the data science minor and this is my final cognitive science elective to fulfill the secondary major requirements. I have also done some work in R for Dr. Andrew Dacks in the biology department using the Drosophila Connectome database. My interests lie in the use of data science to understand more about cognition and the brain. I think technology is a great tool for analyzing data to make conclusions about human cognition. This course will allow me to apply both my background in cognitive science and my skills acquired from data science courses to learn more about why humans act the way we do.
You said:
Hello! Could you please write me some code in R that reads in a dataframe, mutates a column and then creates a plot of that data.
ChatGPT said:
Here’s an example of R code that demonstrates how to read in a dataframe, mutate a column, and then create a plot. We’ll use the ggplot2 package for plotting and the dplyr package for data manipulation.
Example Code:
# Load necessary libraries
library(dplyr)
library(ggplot2)
# Create a sample dataframe
data <- data.frame(
ID = 1:10,
Value = c(5, 10, 15, 20, 25, 30, 35, 40, 45, 50)
)
# Print original dataframe
print("Original Data:")
print(data)
# Mutate a column: Add a new column 'Value_Squared' (square of the 'Value' column)
data <- data %>%
mutate(Value_Squared = Value^2)
# Print mutated dataframe
print("Mutated Data:")
print(data)
# Create a plot: Scatterplot of 'ID' vs 'Value_Squared'
plot <- ggplot(data, aes(x = ID, y = Value_Squared)) +
geom_point(color = "blue", size = 3) +
geom_line(color = "darkblue") +
labs(
title = "Scatterplot of ID vs Value_Squared",
x = "ID",
y = "Value Squared"
) +
theme_minimal()
# Display the plot
print(plot)
Explanation:
Let me know if you’d like further modifications or explanations!
You said:
Hello! What can you tell me about cognition and computation?
ChatGPT said:
Cognition and computation are closely linked fields that explore how mental processes (such as thinking, learning, memory, and perception) can be understood and modeled, often using the framework of information processing. Here’s an overview of each and how they intersect:
Cognition
Cognitive science integrates insights from psychology, neuroscience, linguistics, philosophy, and artificial intelligence (AI) to study these processes.
Computation
Intersection of Cognition and Computation
Applications
If you’d like to explore a specific aspect, feel free to ask!
You said:
Could you please tell me more about how large language models employ cognitive science principles?
ChatGPT said:
Large language models (LLMs), such as GPT, leverage several principles from cognitive science, even if indirectly. These principles inform how these models are designed, trained, and used to simulate aspects of human-like cognition. Let’s break down how LLMs relate to key cognitive science concepts:
Challenges and Limitations
Conclusion
You said:
How can computational models be used to learn more about higher cognitive functions?
ChatGPT said:
Computational models are powerful tools for studying higher cognitive functions—such as reasoning, decision-making, problem-solving, and language comprehension—by providing a framework to simulate, analyze, and test theories about how these processes work. Here’s how they contribute to advancing our understanding of higher cognition:
Examples of Computational Models for Higher Cognitive Functions
Challenges and Limitations
Conclusion
Computational models are invaluable for exploring higher cognitive functions, offering a systematic way to test hypotheses, reveal mechanisms, and integrate findings across disciplines. As these models grow more sophisticated, they will continue to bridge gaps between theory, data, and real-world applications.
I am very interested in the intersection of cognition and computation because it combines my interests of cognitive science and data science. I have a special interest in cognitive neuroscience and find it fascinating that technology can be used to model this. One topic in cognition and computation that interests me is Theory of Mind. Last semester, I read a recent publication that evaluated the ability of Large Language Models to complete Theory of Mind tasks. I am interested in how Theory of Mind is modeled and if this phenomenon can be accurately replicated in machines.
I am also interested in the role the computational model plays in theories of embodied cognition. I am intrigued by how computation can be used as a tool when looking at cognition from a 4E approach. Since computers do not have a mind or a body, how can they be used to model cognitive processes that require extra bodily interactions? Are there other uses for data science in cognitive science that are not centered around modeling cognitive processes?
Based on my previous research, Theory of Mind is an embodied and embedded phenomenon that requires extra bodily processes. Because of this, I wonder how computation can be used to learn about Theory of Mind in a relevant way.