Applied Data project 2 - Introduction to Statistics

Relation between Happiness Score and GDP per Capita

Svara Rahul Masurekar(4027297) and Spandana Balnad Kattamane(4022130)

Last updated: 15 October, 2023

Introduction

Problem Statement

In this analysis, we seek to investigate the relationship between a country’s GDP per capita and its happiness score in the year [year]. We aim to determine if economic prosperity, as measured by GDP per capita, significantly impacts the happiness levels of different countries. Understanding this relationship can have implications for public policy and decision-making. We will utilize open data sources, including the World Happiness Report and World Bank data, to explore this question.

Data

# Data Cont.

Variables of World Happiness data:

Descriptive Statistics and Visualisation

final_data %>% 
  group_by(Continent) %>% 
  summarise(
    Mean = mean(`World Population Percentage`, na.rm = TRUE),
    Median = median(`World Population Percentage`, na.rm = TRUE),
    StandardDeviation = sd(`World Population Percentage`, na.rm = TRUE),
    FirstQuartile = quantile(`World Population Percentage`, probs = 0.25, na.rm = TRUE),
    ThirdQuartile = quantile(`World Population Percentage`, probs = 0.75, na.rm = TRUE),
    Min = min(`World Population Percentage`, na.rm = TRUE),
    Max = max(`World Population Percentage`, na.rm = TRUE),
    MissingValues = sum(is.na(`World Population Percentage`))
  )
final_data %>% 
  group_by(Continent) %>% 
  summarise(
    Mean = mean(`Density (per km²)`, na.rm = TRUE),
    Median = median(`Density (per km²)`, na.rm = TRUE),
    StandardDeviation = sd(`Density (per km²)`, na.rm = TRUE),
    FirstQuartile = quantile(`Density (per km²)`, probs = 0.25, na.rm = TRUE),
    ThirdQuartile = quantile(`Density (per km²)`, probs = 0.75, na.rm = TRUE),
    Min = min(`Density (per km²)`, na.rm = TRUE),
    Max = max(`Density (per km²)`, na.rm = TRUE),
    MissingValues = sum(is.na(`Density (per km²)`))
  )
final_data$Happiness_Score %>% boxplot(ylab = "Happiness Score") # boxplot

Happiness Score Box Plot:

final_data$GDP_Per_Capita %>% boxplot(ylab = "GDP per Capita") # boxplot

GDP per Capita Box Plot:

granova.ds(data.frame(final_data$Happiness_Score, final_data$GDP_Per_Capita),
           xlab = "Happiness Score",
           ylab = "GDP per Capita"
           )

##             Summary Stats
## n                 120.000
## mean(x)          5645.250
## mean(y)          1443.333
## mean(D=x-y)      4201.917
## SD(D)             801.050
## ES(D)               5.246
## r(x,y)              0.780
## r(x+y,d)            0.879
## LL 95%CI         4057.121
## UL 95%CI         4346.712
## t(D-bar)           57.462
## df.t              119.000
## pval.t              0.000
plot(final_data$Happiness_Score, final_data$GDP_Per_Capita)

Hypothesis Testing

Alternative Hypothesis (H_1): There is a statistically significant correlation between a country’s GDP per capita and its citizens’ happiness. The correlation coefficient (ρ) is not equal to zero.

Let us set the significance value(alpha) = 0.05

After conducting the hypothesis test using t-test for co-relation between GDP of a country and its Happiness Score, we get the p-value as 4.914503e-08, which is less than our significance value. Which indicates a significant difference in happiness scores between the two groups. Thus, we reject the null hypothesis.

model1 <- lm(final_data$Happiness_Score ~ final_data$GDP_Per_Capita, data = final_data)

GDP per Capita:

The data shows a roughly linear relationship in the quantile-quantile plot, indicating that the GDP per Capita might be normally distributed. However, there are deviations at both extremes. There are a few noticeable outliers, especially at the upper quantile, which deviate from the expected normal distribution. *The data in the quantile-quantile plot appears to be mostly linear, suggesting that the Happiness Score might be close to normally distributed. However, deviations exist, especially in the lower quantiles. There are a few outliers in the upper quantile that deviate from what would be expected if the data was perfectly normally distributed.

final_data$Happiness_Score %>% qqPlot(dist = "norm", main = "QQ Plot for Happines_Score") #qqPlot

## [1] 120 119
final_data$GDP_Per_Capita %>% qqPlot(dist = "norm", main = "QQ Plot for GDP_per_Capita") #qqPlot

## [1] 92 86
threshold <- 1000

final_data <- final_data %>% mutate(High_GDP_Countries = ifelse(final_data$GDP_Per_Capita >= threshold, "High GDP", "Low GDP"))
final_data
# Perform the t-test
t_test_result <- t.test(final_data$Happiness_Score ~ final_data$High_GDP_Countries)

p_value <- t_test_result$p.value
p_value
## [1] 4.914503e-08

Discussion

Major Findings :

Strengths & Limitations:

Conclusion :

There is strong statistical evidence to reject the null hypothesis. Therefore, it can be concluded that there is a significant difference in happiness scores between countries with GDP per capita above the chosen threshold and those with GDP per capita below the threshold.

References