How this homework works: You will inspect a real survey dataset manually first, then ask an AI tool (Claude, ChatGPT, or Gemini) to inspect the same data, then compare results. Some code is provided for the manual inspection. The cleaning code in Part 3 comes from AI — your job is to run it, evaluate it, and fix what does not work.
Grading reminder: Full sentences. No raw R variable names. Document every AI prompt you used. Code appendix must be present.
library(tidyverse)
library(readxl)
Download CAB_Wave12_Kazakhstan.xlsx from Moodle and
place it in the same folder as this .Rmd file. This dataset
is the Central Asian Barometer (CAB) Wave 12, collected in Kazakhstan in
December 2022 — shortly after Russia’s full-scale invasion of Ukraine in
February 2022.
Research question: Among Kazakhstanis, do those who primarily get news from social media hold more negative views of Russia’s regional influence, compared to those who rely on national state television, after controlling for ethnicity, age group, gender, and income?
cab <- read_excel("/Users/aya/Downloads/CAB_Wave12_Kazakhstan.xlsx")
Run all of these before using any AI tool. Write your observations below each chunk.
dim(cab)
## [1] 1516 28
names(cab)
## [1] "interviewer_gender" "region"
## [3] "urban_rural" "interview_duration"
## [5] "contact_attempts" "primary_news_source"
## [7] "secondary_news_source" "news_frequency"
## [9] "econ_direction_a" "econ_direction_b"
## [11] "country_direction" "view_russia"
## [13] "view_china" "view_usa"
## [15] "russia_favorability" "most_influential_country"
## [17] "ukraine_war_justified" "ukraine_war_impact_kaz"
## [19] "pres_approval" "gender"
## [21] "age" "marital_status"
## [23] "education" "income_bracket"
## [25] "ethnicity" "age_group"
## [27] "weight_v1" "weight_v2"
Write here: How many rows and columns? What does each row represent?
head(cab, 10)
## # A tibble: 10 × 28
## interviewer_gender region urban_rural interview_duration contact_attempts
## <chr> <chr> <chr> <chr> <chr>
## 1 Female Karaganda Urban 00:25:57 First call
## 2 Female Akmola Rural 00:21:21 First call
## 3 Male Almaty Urban 00:24:32 Third call
## 4 Female Astana Ci… Urban 00:49:43 Second call
## 5 Female Almaty Urban 00:19:11 Third call
## 6 Male Kyzylorda Urban 00:16:42 Second call
## 7 Male Almaty Urban 00:16:55 Third call
## 8 Female Turkistan… Rural 00:26:01 Second call
## 9 Female Almaty Urban 00:13:32 Third call
## 10 Male Almaty Ci… Urban 00:11:28 Third call
## # ℹ 23 more variables: primary_news_source <chr>, secondary_news_source <chr>,
## # news_frequency <chr>, econ_direction_a <chr>, econ_direction_b <chr>,
## # country_direction <chr>, view_russia <chr>, view_china <chr>,
## # view_usa <chr>, russia_favorability <chr>, most_influential_country <chr>,
## # ukraine_war_justified <chr>, ukraine_war_impact_kaz <chr>,
## # pres_approval <chr>, gender <chr>, age <dbl>, marital_status <chr>,
## # education <chr>, income_bracket <chr>, ethnicity <chr>, age_group <chr>, …
Write here: Note anything unusual you see in the first 10 rows. Look carefully at the text values — not just the numbers.
summary(cab)
## interviewer_gender region urban_rural interview_duration
## Length :1516 Length :1516 Length :1516 Length :1516
## N.unique : 2 N.unique : 20 N.unique : 3 N.unique : 835
## N.blank : 0 N.blank : 0 N.blank : 0 N.blank : 0
## Min.nchar: 4 Min.nchar: 6 Min.nchar: 5 Min.nchar: 8
## Max.nchar: 6 Max.nchar: 28 Max.nchar: 17 Max.nchar: 8
##
## contact_attempts primary_news_source secondary_news_source news_frequency
## Length :1516 Length :1516 Length :1516 Length :1516
## N.unique : 4 N.unique : 13 N.unique : 13 N.unique : 6
## N.blank : 0 N.blank : 0 N.blank : 0 N.blank : 0
## Min.nchar: 3 Min.nchar: 8 Min.nchar: 8 Min.nchar: 5
## Max.nchar: 11 Max.nchar: 50 Max.nchar: 42 Max.nchar: 21
##
## econ_direction_a econ_direction_b country_direction view_russia
## Length :1516 Length :1516 Length :1516 Length :1516
## N.unique : 7 N.unique : 7 N.unique : 4 N.unique : 6
## N.blank : 0 N.blank : 0 N.blank : 0 N.blank : 0
## Min.nchar: 10 Min.nchar: 10 Min.nchar: 14 Min.nchar: 14
## Max.nchar: 17 Max.nchar: 17 Max.nchar: 17 Max.nchar: 20
##
## view_china view_usa russia_favorability most_influential_country
## Length :1516 Length :1516 Length :1516 Length :1516
## N.unique : 6 N.unique : 6 N.unique : 6 N.unique : 9
## N.blank : 0 N.blank : 0 N.blank : 0 N.blank : 0
## Min.nchar: 14 Min.nchar: 14 Min.nchar: 14 Min.nchar: 4
## Max.nchar: 20 Max.nchar: 20 Max.nchar: 20 Max.nchar: 23
##
## ukraine_war_justified ukraine_war_impact_kaz pres_approval gender
## Length :1516 Length :1516 Length :1516 Length :1516
## N.unique : 6 N.unique : 7 N.unique : 6 N.unique : 2
## N.blank : 0 N.blank : 0 N.blank : 0 N.blank : 0
## Min.nchar: 14 Min.nchar: 14 Min.nchar: 14 Min.nchar: 4
## Max.nchar: 22 Max.nchar: 24 Max.nchar: 19 Max.nchar: 6
##
## age marital_status education income_bracket
## Min. : 18.00 Length :1516 Length :1516 Length :1516
## 1st Qu.: 27.00 N.unique : 5 N.unique : 9 N.unique : 12
## Median : 35.00 N.blank : 0 N.blank : 0 N.blank : 0
## Mean : 42.66 Min.nchar: 6 Min.nchar: 12 Min.nchar: 14
## 3rd Qu.: 47.00 Max.nchar: 19 Max.nchar: 53 Max.nchar: 23
## Max. :999.00
## ethnicity age_group weight_v1 weight_v2
## Length :1516 Length :1516 Min. :0.0296 Min. :0.07632
## N.unique : 3 N.unique : 5 1st Qu.:0.4663 1st Qu.:0.46685
## N.blank : 0 N.blank : 0 Median :0.7790 Median :0.75753
## Min.nchar: 5 Min.nchar: 3 Mean :0.9996 Mean :0.99960
## Max.nchar: 7 Max.nchar: 5 3rd Qu.:1.2842 3rd Qu.:1.27955
## Max. :7.3648 Max. :3.92409
names(cab)
## [1] "interviewer_gender" "region"
## [3] "urban_rural" "interview_duration"
## [5] "contact_attempts" "primary_news_source"
## [7] "secondary_news_source" "news_frequency"
## [9] "econ_direction_a" "econ_direction_b"
## [11] "country_direction" "view_russia"
## [13] "view_china" "view_usa"
## [15] "russia_favorability" "most_influential_country"
## [17] "ukraine_war_justified" "ukraine_war_impact_kaz"
## [19] "pres_approval" "gender"
## [21] "age" "marital_status"
## [23] "education" "income_bracket"
## [25] "ethnicity" "age_group"
## [27] "weight_v1" "weight_v2"
# Look closely at the variables relevant to our research question.
# E1a = view of Russia (Very Unfavorable → Very Favorable)
# A1a = primary news source
# EthnicGrp = ethnicity
# AgeGrp = age group
# DD1 = gender
# DD8 = income bracket
# FinalWgt1, FinalWgt2 = two survey weight variables
cab %>%
select(view_russia, primary_news_source, ethnicity, age_group,
gender, income_bracket, weight_v1, weight_v2) %>%
summary()
## view_russia primary_news_source ethnicity age_group
## Length :1516 Length :1516 Length :1516 Length :1516
## N.unique : 6 N.unique : 13 N.unique : 3 N.unique : 5
## N.blank : 0 N.blank : 0 N.blank : 0 N.blank : 0
## Min.nchar: 14 Min.nchar: 8 Min.nchar: 5 Min.nchar: 3
## Max.nchar: 20 Max.nchar: 50 Max.nchar: 7 Max.nchar: 5
##
## gender income_bracket weight_v1 weight_v2
## Length :1516 Length :1516 Min. :0.0296 Min. :0.07632
## N.unique : 2 N.unique : 12 1st Qu.:0.4663 1st Qu.:0.46685
## N.blank : 0 N.blank : 0 Median :0.7790 Median :0.75753
## Min.nchar: 4 Min.nchar: 14 Mean :0.9996 Mean :0.99960
## Max.nchar: 6 Max.nchar: 23 3rd Qu.:1.2842 3rd Qu.:1.27955
## Max. :7.3648 Max. :3.92409
# Check the unique values in each categorical variable
cat("View of Russia:\n"); print(sort(unique(cab$view_russia)))
## View of Russia:
## [1] "Don't Know (vol.)" "Refused (vol.)" "Somewhat Favorable"
## [4] "Somewhat Unfavorable" "Very Favorable" "Very Unfavorable"
cat("\nPrimary news source:\n"); print(sort(unique(cab$primary_news_source)))
##
## Primary news source:
## [1] "Don't Know (vol.)"
## [2] "Family, friends, or neighbors"
## [3] "Internet"
## [4] "Internet: blogs, others"
## [5] "Local radio"
## [6] "Non-government newspapers published in our country"
## [7] "None / Not interested in news (vol.)"
## [8] "Other foreign news sources"
## [9] "Our country's national television stations"
## [10] "Our government's newspapers"
## [11] "Refused (vol.)"
## [12] "Russian television stations"
## [13] "Social media sites or messengers"
cat("\nEthnicity:\n"); print(sort(unique(cab$ethnicity)))
##
## Ethnicity:
## [1] "Kazakh" "Other" "Russian"
cat("\nAge group:\n"); print(sort(unique(cab$age_group)))
##
## Age group:
## [1] "18-29" "30-39" "40-49" "50-59" "60+"
cat("\nGender:\n"); print(sort(unique(cab$gender)))
##
## Gender:
## [1] "Female" "Male"
cat("\nIncome bracket:\n"); print(sort(unique(cab$income_bracket)))
##
## Income bracket:
## [1] "100,001 - 150,000 tenge" "150,001 - 250,000 tenge"
## [3] "250,001 - 300,000 tenge" "30,001 - 40,000 tenge"
## [5] "40,001 - 50,000 tenge" "50,001 - 60,000 tenge"
## [7] "60,001 - 80,000 tenge" "80,001 - 100,000 tenge"
## [9] "Don't Know (vol.)" "Less than 30,001 tenge"
## [11] "More than 300,000 tenge" "Refused (vol.)"
# The dataset has two weight variables. Compare them.
# The dataset has two weight variables. Compare them.
summary(cab$weight_v1)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 0.0296 0.4663 0.7790 0.9996 1.2842 7.3648
summary(cab$weight_v2)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 0.07632 0.46685 0.75753 0.99960 1.27955 3.92409
cat("Correlation between the two weights:",
round(cor(cab$weight_v1, cab$weight_v2, use = "complete.obs"), 4), "\n")
## Correlation between the two weights: 0.9478
# Check for hidden text characters in string variables
# (Some survey export files contain non-breaking spaces \xa0)
cat("Primary news source values with whitespace issues:\n")
## Primary news source values with whitespace issues:
cab %>%
count(primary_news_source) %>%
filter(grepl("\\s{2,}|^\\s|\\s$", primary_news_source)) %>%
print()
## # A tibble: 0 × 2
## # ℹ 2 variables: primary_news_source <chr>, n <int>
cat("\nView of Russia: check for capitalization inconsistencies:\n")
##
## View of Russia: check for capitalization inconsistencies:
cab %>% count(view_russia) %>% print()
## # A tibble: 6 × 2
## view_russia n
## <chr> <int>
## 1 Don't Know (vol.) 164
## 2 Refused (vol.) 31
## 3 Somewhat Favorable 612
## 4 Somewhat Unfavorable 253
## 5 Very Favorable 216
## 6 Very Unfavorable 240
# Survey data uses text codes ("Don't Know (vol.)", "Refused (vol.)", "Not asked")
# instead of NA. Count how often each appears in our key variables.
special_codes <- c("Don't Know (vol.)", "Don't know (vol.)",
"Refused (vol.)", "Not asked", "Not Asked")
cab %>%
select(view_russia, primary_news_source, ethnicity, gender, income_bracket) %>%
summarise(across(everything(),
~ sum(. %in% special_codes, na.rm = TRUE),
.names = "n_special_{.col}"))
## # A tibble: 1 × 5
## n_special_view_russia n_special_primary_news_source n_special_ethnicity
## <int> <int> <int>
## 1 195 36 0
## # ℹ 2 more variables: n_special_gender <int>, n_special_income_bracket <int>
Write here (Question 1): List every data quality problem you found. For each problem: what is the issue, which variable(s) are affected, and why is it a problem for your analysis? Write this as numbered points with 2–3 sentences each.
💡 What to look for in survey data: inconsistent capitalization (same value written differently), non-breaking space characters () mixed into text, “Don’t Know” and “Refused” coded as text strings rather than NA, duplicate-looking columns that appear to ask the same question, two versions of a weighting variable with no documentation about which to use, and columns where most respondents got “Not asked” because of a skip pattern.
Upload CAB_Wave12_Kazakhstan.xlsx to your AI tool and
use exactly this prompt:
“This file is the Central Asian Barometer survey, Wave 12, collected in Kazakhstan in December 2022. I am trying to study whether primary news source (social media vs. state television) predicts views of Russia, controlling for ethnicity, age, gender, and income. The key variables are: E1a (view of Russia), A1a (primary news source), EthnicGrp, AgeGrp, DD1, DD8, FinalWgt1, FinalWgt2. Please inspect this dataset and identify all data quality issues relevant to this analysis. Look specifically for: inconsistent text values, hidden characters, missing-value codes stored as text strings, duplicate or redundant columns, and any survey-specific concerns.”
Question 2 — write your answer here:
List the issues the AI identified. For each one, note whether you had already found it yourself. Then write 3–4 sentences comparing your manual inspection to the AI’s output: was it more thorough, less thorough, or about the same? Did the AI recognize that “Don’t Know (vol.)” is a meaningful survey response, not just a data error? What does this tell you about the limits of AI for survey data review?
Paste the AI’s issue list here (verbatim):
[Paste the AI's response here]
Use this follow-up prompt in the same AI session:
“Now write R code to clean the E1a, A1a, EthnicGrp, AgeGrp, DD1, and DD8 columns for analysis. Specifically: (1) replace ‘Don’t Know (vol.)’, ‘Refused (vol.)’, ‘Not asked’, and any variant of these with NA; (2) standardize capitalization and remove leading/trailing whitespace; (3) create a numeric version of E1a (Very Unfavorable=1, Somewhat Unfavorable=2, Somewhat Favorable=3, Very Favorable=4); (4) create a binary variable for news source: 1 = social media (any response mentioning ‘social media’ or ‘messengers’), 0 = national TV. After cleaning, show the frequency tables for all cleaned variables.”
Question 3 — paste and run the AI’s cleaning code below.
Edit the chunk as needed to make it run. If the AI gave code in sections, combine it here.
# PASTE THE AI'S CLEANING CODE HERE
# Run it and fix any errors
library(readxl)
library(dplyr)
library(stringr)
# Load data
df <- read_excel("CAB_Wave12_Kazakhstan.xlsx")
# Function to clean text variables
clean_text <- function(x) {
x <- as.character(x)
x <- str_squish(x) # remove extra/leading/trailing whitespace
x <- str_to_lower(x) # standardize capitalization
# Convert non-substantive responses to NA
x[x %in% c(
"don't know (vol.)",
"refused (vol.)",
"not asked",
"don't know",
"refused"
)] <- NA_character_
x
}
# Clean the six variables
df <- df %>%
mutate(
view_russia_clean = clean_text(view_russia),
primary_news_source_clean = clean_text(primary_news_source),
ethnicity_clean = clean_text(ethnicity),
age_group_clean = clean_text(age_group),
gender_clean = clean_text(gender),
income_bracket_clean = clean_text(income_bracket)
)
# Numeric version of E1a
df <- df %>%
mutate(
view_russia_numeric = case_when(
view_russia_clean == "very unfavorable" ~ 1,
view_russia_clean == "somewhat unfavorable" ~ 2,
view_russia_clean == "somewhat favorable" ~ 3,
view_russia_clean == "very favorable" ~ 4,
TRUE ~ NA_real_
)
)
# Binary news-source variable
# 1 = social media / messengers
# 0 = national television
# All other responses = NA
df <- df %>%
mutate(
news_source_binary = case_when(
str_detect(primary_news_source_clean, "social media|messengers") ~ 1,
primary_news_source_clean == "our country's national television stations" ~ 0,
TRUE ~ NA_real_
)
)
# Frequency tables
table(df$view_russia_clean, useNA = "ifany")
##
## somewhat favorable somewhat unfavorable very favorable
## 612 253 216
## very unfavorable <NA>
## 240 195
table(df$view_russia_numeric, useNA = "ifany")
##
## 1 2 3 4 <NA>
## 240 253 612 216 195
table(df$primary_news_source_clean, useNA = "ifany")
##
## family, friends, or neighbors
## 20
## internet
## 250
## internet: blogs, others
## 138
## local radio
## 5
## non-government newspapers published in our country
## 1
## none / not interested in news (vol.)
## 73
## other foreign news sources
## 1
## our country's national television stations
## 283
## our government's newspapers
## 7
## russian television stations
## 26
## social media sites or messengers
## 676
## <NA>
## 36
table(df$news_source_binary, useNA = "ifany")
##
## 0 1 <NA>
## 283 676 557
table(df$ethnicity_clean, useNA = "ifany")
##
## kazakh other russian
## 977 252 287
table(df$age_group_clean, useNA = "ifany")
##
## 18-29 30-39 40-49 50-59 60+
## 475 466 269 167 139
table(df$gender_clean, useNA = "ifany")
##
## female male
## 623 893
table(df$income_bracket_clean, useNA = "ifany")
##
## 100,001 - 150,000 tenge 150,001 - 250,000 tenge 250,001 - 300,000 tenge
## 192 278 172
## 30,001 - 40,000 tenge 40,001 - 50,000 tenge 50,001 - 60,000 tenge
## 11 15 25
## 60,001 - 80,000 tenge 80,001 - 100,000 tenge less than 30,001 tenge
## 48 89 7
## more than 300,000 tenge <NA>
## 462 217
Write here: Did the code run without errors? If not, what errors did you get and how did you fix them? In 3–4 sentences, describe any corrections you made to the AI’s code. Pay special attention to whether the AI correctly handled the non-breaking space characters () and whether it handled both capitalization variants of “Don’t Know (vol.)”.
# Compare raw vs cleaned values for view of Russia
cat("Raw view of Russia distribution:\n")
## Raw view of Russia distribution:
table(df$view_russia, useNA = "always")
##
## Don't Know (vol.) Refused (vol.) Somewhat Favorable
## 164 31 612
## Somewhat Unfavorable Very Favorable Very Unfavorable
## 253 216 240
## <NA>
## 0
cat("\nCleaned numeric view of Russia distribution:\n")
##
## Cleaned numeric view of Russia distribution:
if ("view_russia_numeric" %in% names(df)) {
table(df$view_russia_numeric, useNA = "always")
} else {
cat("Variable view_russia_numeric not found — check AI's variable naming\n")
}
##
## 1 2 3 4 <NA>
## 240 253 612 216 195
cat("\nCleaned news source (1 = social media, 0 = national TV) distribution:\n")
##
## Cleaned news source (1 = social media, 0 = national TV) distribution:
if ("news_source_binary" %in% names(df)) {
table(df$news_source_binary, useNA = "always")
} else {
cat("Variable news_source_binary not found — check AI's variable naming\n")
}
##
## 0 1 <NA>
## 283 676 557
Question 4 — write your answer here:
Report how many valid observations (non-NA) remain in E1a_num after cleaning. Describe how removing “Don’t Know” and “Refused” responses changes the sample. In 2–3 sentences, explain whether this non-response is likely to be random or systematic — and what that means for who your cleaned analysis actually represents.
Question 5 — write your answer here:
There are cleaning decisions the AI cannot make. Describe two such decisions you faced. What did you decide, and why? Write 2–3 sentences per decision.
💡 Examples specific to this data: Which weight variable (FinalWgt1 vs. FinalWgt2) should you use — and does it matter? How should you handle respondents who said “Don’t Know” on the Russia views question — drop them, or keep them as a separate category? Should you group social media platforms (Instagram, TikTok, WhatsApp) into one category or treat them separately? AI can generate code but cannot answer these questions without knowing the theory behind your research question.
Question 6 — write 2 paragraphs:
Paragraph 1: Where is AI most useful in the data cleaning workflow for survey data specifically, and why?
Paragraph 2: Where does AI fall short with survey data compared to a simple spreadsheet of country statistics? Name at least one type of error that AI commonly makes when cleaning survey data — and explain why it happens.
Required: List every prompt you used with the AI in this homework. You do not need to paste the full responses, but list the prompts in order.
Knit → HTML. Submit the .html and .Rmd to
Moodle.