Siri Chandana Chukka
2024-12-06
“In this project, I aim to analyze the global issue of violence against women (VaW), particularly focusing on the factors that influence the frequency and patterns of violent assaults, with a focus on serious assaults.”
“VaW is a serious global problem that impacts women’s physical, emotional, and mental well-being. It also leads to economic consequences due to the departure of women from the workforce and costs related to healthcare and legal services.”
“How do violent assaults and other forms of violence against women differ by location and time of year?”
“What is the relationship between the frequency of serious attacks and historical patterns of violence?”
“The dataset I used comes from the United Nations Office on Drugs and Crime (UNODC), which provides comprehensive data on violent crimes, including sexual violence and assaults, broken down by country, region, and year.”
“I imported the dataset and adjusted column names for clarity.”
“I removed unnecessary columns, omitted missing values, and converted relevant variables (like ‘Year’ and ‘VALUE’) to numeric format.”
“I filtered the data to include only records from 2004 onwards, and categorized the severity of violence into ‘Low’, ‘Medium’, and ‘High’ based on the number of reported incidents.”
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.1.4 ✔ readr 2.1.5
## ✔ forcats 1.0.0 ✔ stringr 1.5.1
## ✔ ggplot2 3.5.1 ✔ tibble 3.2.1
## ✔ lubridate 1.9.3 ✔ tidyr 1.3.1
## ✔ purrr 1.0.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
# loading dataset separately in an R console or script
library(readr)
data_cts_violent_and_sexual_crime <- read_csv("C:/Users/siric/Downloads/data_cts_violent_and_sexual_crime.csv")## New names:
## • `` -> `...3`
## • `` -> `...4`
## • `` -> `...5`
## • `` -> `...6`
## • `` -> `...7`
## • `` -> `...8`
## • `` -> `...9`
## • `` -> `...10`
## • `` -> `...11`
## • `` -> `...12`
## • `` -> `...13`
## Warning: One or more parsing issues, call `problems()` on your data frame for details,
## e.g.:
## dat <- vroom(...)
## problems(dat)
## Rows: 26116 Columns: 13
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (11): UNODC, unodc_ddds@un.org, ...3, ...4, ...5, ...6, ...7, ...8, ...9...
## dbl (2): ...10, ...12
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
## # A tibble: 6 × 13
## UNODC `unodc_ddds@un.org` ...3 ...4 ...5 ...6 ...7 ...8 ...9 ...10
## <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <dbl>
## 1 16/05/2024 <NA> <NA> <NA> <NA> <NA> <NA> <NA> <NA> NA
## 2 Iso3_code Country Regi… Subr… Indi… Dime… Cate… Sex Age NA
## 3 AZE Azerbaijan Asia West… Viol… by t… Seri… Total Total 2003
## 4 BEL Belgium Euro… West… Viol… by t… Seri… Total Total 2003
## 5 BGR Bulgaria Euro… East… Viol… by t… Seri… Total Total 2003
## 6 BHR Bahrain Asia West… Viol… by t… Seri… Total Total 2003
## # ℹ 3 more variables: ...11 <chr>, ...12 <dbl>, ...13 <chr>
colnames(data_cts_violent_and_sexual_crime) <- c("Iso3_code", "Country", "Region", "Subregion", "indicator","Dimension", "Category", "Sex", "Age", "Year","Unit", "VALUE", "Source")# Data Cleaning and Filtering
data_cleaned <- data_cts_violent_and_sexual_crime %>%
select(-Unit) %>%
na.omit() %>%
mutate(Year = as.numeric(Year), VALUE = as.numeric(VALUE))
# Filtering data for years >= 2004
data_filtered <- data_cleaned %>%
filter(Year >= 2004)
# Create a Violence Level categorization
data_filtered <- data_filtered %>%
mutate(Violence_Level = case_when(
VALUE > 10000 ~ "High",
VALUE > 1000 ~ "Medium",
TRUE ~ "Low"
))data_aggregated <- data_filtered %>%
filter(Category == "Serious assault") %>%
group_by(Region, Year) %>%
summarise(Total_Violence = sum(VALUE, na.rm = TRUE))## `summarise()` has grouped output by 'Region'. You can override using the
## `.groups` argument.
## # A tibble: 6 × 3
## # Groups: Region [1]
## Region Year Total_Violence
## <chr> <dbl> <dbl>
## 1 Africa 2004 209266.
## 2 Africa 2005 212124.
## 3 Africa 2006 212770.
## 4 Africa 2007 232622.
## 5 Africa 2008 215769.
## 6 Africa 2009 147930.
“I used a line graph to visualize how violent assaults evolved over time, categorized by region.”
library(ggplot2)
ggplot(data_aggregated, aes(x = Year, y = Total_Violence, color = Region)) +
geom_line() +
theme_minimal() +
labs(title = "Total Serious Assaults Over Time by Region",
x = "Year",
y = "Total Serious Assaults")“The line graph reveals that the total number of serious assaults has fluctuated over the years, with regional differences becoming more apparent in certain years.”
“There is a noticeable difference in the frequency of serious assaults based on regional context and time of year.”
“This analysis highlights how violence against women varies by region and year, and how historical events and societal changes impact these trends.”
“Understanding these patterns is critical for policymakers and organizations aiming to reduce violence against women. It can help target interventions in areas where the violence rate is high or increasing.”