Svara Rahul Masurekar(4027297) and Spandana Balnad Kattamane(4022130)
Last updated: 15 October, 2023
It is not necessary (That is, it is optional and not compulsory) but if you like you can publish your presentation to RPubs (see here) and add this link to your presentation here.
Rpubs link comes here: www………
In an increasingly globalized world, nations aspire not just for economic growth, but also for the well-being and happiness of their citizens. The World Happiness Index provides an invaluable tool, quantifying happiness and its various determinants across the globe.
Beyond the aforementioned determinants, demographic attributes of countries, such as their population size, density, growth rate, and continental location, play a potentially pivotal role. The intersection of these demographic details with happiness indicators is a less explored, yet crucial area of research.
The components contributing to happiness are multifaceted. They range from economic factors such as GDP per capita, to societal elements like social support, and even individual rights such as freedom to make life choices.
By bridging the correlation between happiness and demographic factors, we aspire to enlighten policy makers, social scientists, and global leaders on pathways to elevate societal well-being.
In this analysis, we seek to investigate the relationship between a country’s GDP per capita and its happiness score in the year [year]. We aim to determine if economic prosperity, as measured by GDP per capita, significantly impacts the happiness levels of different countries. Understanding this relationship can have implications for public policy and decision-making. We will utilize open data sources, including the World Happiness Report and World Bank data, to explore this question.
Use of Statistics to Solve the Problem:
Hypothesis Testing:
Visual Interpretation: Through scatter plots and other graphical representations, we visually inspect the distribution and relationship between GDP per capita and happiness scores. These visual aids complement our statistical tests by providing a more intuitive understanding of data trends and patterns.
Correlation Analysis: Beyond the t-test, we can also calculate the correlation coefficient between GDP per capita and happiness scores. This coefficient will provide a measure of the strength and direction of the linear relationship between these two variables, further enhancing our analysis.
Categorical Analysis: By segmenting countries into “High GDP” and “Low GDP” based on a defined GDP per capita threshold, we can draw comparisons between these two groups. This allows for a clearer understanding of how GDP per capita categories might relate to happiness score.
The links to our datasets, World happiness report and World population, are provided below. We received our datasets from the Kaggle website.
https://www.kaggle.com/datasets/iamsouravbanerjee/world-population-dataset
https://www.kaggle.com/datasets/mathurinache/world-happiness-report
Variables of World population dataset :
CCA3: This is likely a three-letter country code. It’s a unique identifier for each country.
Country/Territory: This variable represents the name of the country or territory.
Capital: This is the capital city of the respective country or territory.
Continent: The continent where the country or territory is located.
year Population: The estimated population of the country or territory in that specific year .
Area (km²): The total land area of the country or territory, measured in square kilometers.
Density (per km²): The population density, calculated as the population divided by the land area. It represents how many people live in each square kilometer of the country or territory.
Growth Rate: This variable might indicate the rate at which the population of the country or territory is growing or declining.
World Population Percentage: The percentage of the world’s total population that resides in the respective country or territory.
# Data Cont.
Variables of World Happiness data:
RANK: A numerical ranking based on the happiness score, with lower numbers indicating higher happiness.
Country: The name of the country for which the data pertains.
Happiness Score: A score that quantifies the happiness level of residents in a particular country.
Whisker-high & Whisker-low: These represent the upper and lower bounds, respectively, of the confidence interval around the happiness score. They give a range within which the true happiness score is likely to fall.
Explained by: GDP per capita: The extent to which Gross Domestic Product per capita contributes to the happiness score.
Explained by: Social support: The contribution of social support networks (like family and friends) to the happiness score.
Explained by: Perceptions of corruption: The degree to which perceptions of corruption (or lack thereof) in government and business influence the happiness score.
The variable “Continent” of the “final_data” dataframe is converted into a factor variable.
Factor variable : “Continent” that classifies each observation in final dataset into one of the continents of the world. It represents the geographical categorization of the data.
The “Continent” factor has six levels: “Africa”, “Asia”, “Europe”, “North America”, “Oceania”, and “South America”. Each level signifies a specific continent, ensuring clarity in data representation.
World Population Percentage by Continent:
The code groups the dataset by continents and then calculates a
range of statistical metrics for each continent’s
World Population Percentage.
Metrics include mean, median, standard deviation, quartile values, as well as the smallest, largest, and count of missing values.
Density (per km²) by Continent:
Similarly, the data is grouped by continents, focusing on the
Density (per km²) column.
For each continent, it provides metrics such as average density, middle value, variation in density, along with min, max, and missing value counts.
The code groups the dataset by continents and then calculates a
range of statistical metrics for each continent’s
World Population Percentage.
Metrics include mean, median, standard deviation, quartile values, as well as the smallest, largest, and count of missing values.
Similarly, the data is grouped by continents, focusing on the
Density (per km²) column.
For each continent, it provides metrics such as average density, middle value, variation in density, along with min, max, and missing value counts.
final_data %>%
group_by(Continent) %>%
summarise(
Mean = mean(`World Population Percentage`, na.rm = TRUE),
Median = median(`World Population Percentage`, na.rm = TRUE),
StandardDeviation = sd(`World Population Percentage`, na.rm = TRUE),
FirstQuartile = quantile(`World Population Percentage`, probs = 0.25, na.rm = TRUE),
ThirdQuartile = quantile(`World Population Percentage`, probs = 0.75, na.rm = TRUE),
Min = min(`World Population Percentage`, na.rm = TRUE),
Max = max(`World Population Percentage`, na.rm = TRUE),
MissingValues = sum(is.na(`World Population Percentage`))
)final_data %>%
group_by(Continent) %>%
summarise(
Mean = mean(`Density (per km²)`, na.rm = TRUE),
Median = median(`Density (per km²)`, na.rm = TRUE),
StandardDeviation = sd(`Density (per km²)`, na.rm = TRUE),
FirstQuartile = quantile(`Density (per km²)`, probs = 0.25, na.rm = TRUE),
ThirdQuartile = quantile(`Density (per km²)`, probs = 0.75, na.rm = TRUE),
Min = min(`Density (per km²)`, na.rm = TRUE),
Max = max(`Density (per km²)`, na.rm = TRUE),
MissingValues = sum(is.na(`Density (per km²)`))
)granova.ds(data.frame(final_data$Happiness_Score, final_data$GDP_Per_Capita),
xlab = "Happiness Score",
ylab = "GDP per Capita"
)## Summary Stats
## n 120.000
## mean(x) 5645.250
## mean(y) 1443.333
## mean(D=x-y) 4201.917
## SD(D) 801.050
## ES(D) 5.246
## r(x,y) 0.780
## r(x+y,d) 0.879
## LL 95%CI 4057.121
## UL 95%CI 4346.712
## t(D-bar) 57.462
## df.t 119.000
## pval.t 0.000
Alternative Hypothesis (H_1): There is a statistically significant correlation between a country’s GDP per capita and its citizens’ happiness. The correlation coefficient (ρ) is not equal to zero.
Let us set the significance value(alpha) = 0.05
After conducting the hypothesis test using t-test for co-relation between GDP of a country and its Happiness Score, we get the p-value as 4.914503e-08, which is less than our significance value. Which indicates a significant difference in happiness scores between the two groups. Thus, we reject the null hypothesis.
GDP per Capita:
The data shows a roughly linear relationship in the quantile-quantile plot, indicating that the GDP per Capita might be normally distributed. However, there are deviations at both extremes. There are a few noticeable outliers, especially at the upper quantile, which deviate from the expected normal distribution. *The data in the quantile-quantile plot appears to be mostly linear, suggesting that the Happiness Score might be close to normally distributed. However, deviations exist, especially in the lower quantiles. There are a few outliers in the upper quantile that deviate from what would be expected if the data was perfectly normally distributed.
## [1] 120 119
## [1] 92 86
threshold <- 1000
final_data <- final_data %>% mutate(High_GDP_Countries = ifelse(final_data$GDP_Per_Capita >= threshold, "High GDP", "Low GDP"))
final_data# Perform the t-test
t_test_result <- t.test(final_data$Happiness_Score ~ final_data$High_GDP_Countries)
p_value <- t_test_result$p.value
p_value## [1] 4.914503e-08
Major Findings :
GDP & Happiness Correlation: A positive association exists between GDP per capita and a country’s happiness score; as one rises, so does the other.
Normal Distribution: The majority of data points fit a normal distribution, suggesting consistent trends among most countries.
Presence of Outliers: Deviations at the extremes in Q-Q plots indicate outliers or countries with exceptional happiness scores or GDP values.
Strengths & Limitations:
The data set covers a wide range of countries, giving a comprehensive view of global trends.
The use of Q-Q plots allowed us to visually inspect the normality of our data, aiding in the understanding of its distribution.
Correlation does not imply causation. While there’s a relationship between GDP and happiness, we can’t say definitively that a higher GDP causes more happiness.
Potential outliers might skew the results, leading to potential misinterpretations.
Conclusion :
A positive correlation exists between GDP per capita and happiness score, indicating nations with higher GDP tend to report greater happiness.
The data mostly aligns with a normal distribution, but deviations at extremes highlight potential outliers or exceptionally high-performing countries.
Such outliers suggest varying factors, beyond GDP, might significantly influence a country’s happiness score.
There is strong statistical evidence to reject the null hypothesis. Therefore, it can be concluded that there is a significant difference in happiness scores between countries with GDP per capita above the chosen threshold and those with GDP per capita below the threshold.
ACHÉ, M. (n.d). World Happiness Report up to 2022 [.csv] www.kaggle.com website accessed 09 October 2023. https://www.kaggle.com/datasets/mathurinache/world-happiness-report
www.kaggle.com (n.d.). World Population Dataset [csv] www.kaggle.com website accessed 09 October 2023. https://www.kaggle.com/datasets/iamsouravbanerjee/world-population-dataset