Hypothesis Testing On Adult Smoking Prevalence

Introduction

Sukhvir Singh Chahal (s4019923) and Ruhan Ahmed Malik (s4024339)

Last updated: 15th october 2023

Introduction

The dataset offers a thorough perspective on smoking prevalence in the United States Covering the years from 1984 to 2019 Our analysis centers on two important variables: “GENDER,” which categorizes the data into “Male,” and “Female,” and “PERCENT” reflecting smoking percentages within these gender groups.

Problem Statement

The main objective is to gain insights into the patterns of smoking behavior within the U.S. population and to explore whether there are significant difference in smoking rates between Male and Female Our analysis is centered on two separate t-tests: a two-sample t-test to compare male and female smoking percentages.

Two-Sample t-test for Male and Female Smoking Percentages: a two-sample t-test, we would try to determine whether there exists a statistically significant smoking behavior between these two gender categories.

Data

The Dataset is taken From https://catalog.data.gov/dataset/adult-cigarette-and-tobacco-use-prevalence-31e2f The dataset provides a comprehensive perspective on smoking prevalence and associated confidence intervals, covering the years from 1984 to 2019 in the United states. Within this dataset, we focus our analysis on two key variables:

GENDER: This variable classifies the data into three distinct groups, namely “Male,” and “Female”.

PERCENT: It represents the percentage of individuals within specific gender groups who have reported their smoking behavior.

Data Preprocessing

adult_smoking <- read.csv("adult-smoking-prevalence.csv")
adult_smoking$GENDER<- adult_smoking$GENDER %>% factor(levels=c("Male","Female","Total"))
Male_smoking <- adult_smoking %>% filter(GENDER =="Male")
Female_smoking <- adult_smoking %>% filter(GENDER =="Female")
Adult_smoking_New <- union(Male_smoking,Female_smoking)
str(Adult_smoking_New)
## 'data.frame':    72 obs. of  6 variables:
##  $ YEAR      : int  1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 ...
##  $ COMPARISON: chr  "Definition Expanded in 1996" "Definition Expanded in 1996" "Definition Expanded in 1996" "Definition Expanded in 1996" ...
##  $ GENDER    : Factor w/ 3 levels "Male","Female",..: 1 1 1 1 1 1 1 1 1 1 ...
##  $ PERCENT   : num  26.9 28.2 28 22.8 25.6 22.5 21.1 22.7 22.3 20.7 ...
##  $ LOWER95   : num  22.8 24.6 24.6 19.8 22.9 20 18.8 20.5 20.4 19.3 ...
##  $ UPPER95   : num  31 31.8 31.4 25.8 28.2 25 23.3 25 24.2 22.1 ...
head(Adult_smoking_New)
tail(Adult_smoking_New)

Data Preprocessing cont.

glimpse(Male_smoking)
## Rows: 36
## Columns: 6
## $ YEAR       <int> 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993,…
## $ COMPARISON <chr> "Definition Expanded in 1996", "Definition Expanded in 1996…
## $ GENDER     <fct> Male, Male, Male, Male, Male, Male, Male, Male, Male, Male,…
## $ PERCENT    <dbl> 26.9, 28.2, 28.0, 22.8, 25.6, 22.5, 21.1, 22.7, 22.3, 20.7,…
## $ LOWER95    <dbl> 22.8, 24.6, 24.6, 19.8, 22.9, 20.0, 18.8, 20.5, 20.4, 19.3,…
## $ UPPER95    <dbl> 31.0, 31.8, 31.4, 25.8, 28.2, 25.0, 23.3, 25.0, 24.2, 22.1,…
glimpse(Female_smoking)
## Rows: 36
## Columns: 6
## $ YEAR       <int> 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993,…
## $ COMPARISON <chr> "Definition Expanded in 1996", "Definition Expanded in 1996…
## $ GENDER     <fct> Female, Female, Female, Female, Female, Female, Female, Fem…
## $ PERCENT    <dbl> 23.0, 25.2, 23.3, 20.1, 19.9, 19.8, 17.9, 15.8, 17.8, 15.8,…
## $ LOWER95    <dbl> 19.7, 22.1, 20.5, 17.7, 17.8, 17.7, 15.9, 14.0, 16.2, 14.6,…
## $ UPPER95    <dbl> 26.3, 28.2, 26.0, 22.6, 22.0, 22.0, 19.8, 17.6, 19.4, 16.9,…
glimpse(Adult_smoking_New)
## Rows: 72
## Columns: 6
## $ YEAR       <int> 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993,…
## $ COMPARISON <chr> "Definition Expanded in 1996", "Definition Expanded in 1996…
## $ GENDER     <fct> Male, Male, Male, Male, Male, Male, Male, Male, Male, Male,…
## $ PERCENT    <dbl> 26.9, 28.2, 28.0, 22.8, 25.6, 22.5, 21.1, 22.7, 22.3, 20.7,…
## $ LOWER95    <dbl> 22.8, 24.6, 24.6, 19.8, 22.9, 20.0, 18.8, 20.5, 20.4, 19.3,…
## $ UPPER95    <dbl> 31.0, 31.8, 31.4, 25.8, 28.2, 25.0, 23.3, 25.0, 24.2, 22.1,…

Descriptive Statistics and Visualisation

Male_smoking %>% summarise(Min = min(PERCENT),
                  Q1 = quantile(PERCENT,probs = .25),
                  Median = median(PERCENT),
                  Q3 = quantile(PERCENT,probs = .75), 
                  Max = max(PERCENT),
                  Mean = mean(PERCENT),
                  SD = sd(PERCENT))
Female_smoking %>% summarise(Min = min(PERCENT),
                  Q1 = quantile(PERCENT,probs = .25),
                  Median = median(PERCENT),
                  Q3 = quantile(PERCENT,probs = .75), 
                  Max = max(PERCENT),
                  Mean = mean(PERCENT),
                  SD = sd(PERCENT)) 

(For Two Sample t-test) Through Descriptive Statistics, The mean percent of Male smokers is Not equal to that of Female Group, but we are not sure yet whether it is significantly different or not.

Decsriptive Statistics and Visualistaion Cont. 1

Male_smoking$PERCENT %>% qqPlot(dist="norm")

## [1] 2 3

Decsriptive Statistics and Visualistaion Cont. 2

Male_smoking$PERCENT%>%  
boxplot(main="Male Smoking Percentage", ylab="% Smoking", col = "grey")

Decsriptive Statistics and Visualistaion Cont. 3

Female_smoking$PERCENT %>% qqPlot(dist="norm")

## [1] 2 3

Decsriptive Statistics and Visualistaion Cont. 4

Female_smoking$PERCENT%>%  
boxplot(main="Female Smoking Percentage", ylab="% Smoking", col = "grey")

Levene Test and Explanation

leveneTest(PERCENT ~ GENDER, 
          data = Adult_smoking_New)

According to qq plot and box plots, For both groups, Percent smokers are in normal distribution and moreover Both groups have sample size greater than 30, so the distribution in both groups are normal. As p>0.05, population variances are homogeneous(by levene Test) Thus sampling distribution will approximate a normal distribution and we can apply the two-sample t-test.

Hypothesis Testing

(For Two Sample t-test)

\[H_0: \mu_1 = \mu_2\] \[H_A: \mu_1 \ne \mu_2\]

alpha <- 0.05
Average_Smoker_claim <- 11.5
t1 <- t.test(PERCENT ~ GENDER, 
                data = Adult_smoking_New, 
                var.equal = TRUE, 
                alternative = "two.sided")
t1
## 
##  Two Sample t-test
## 
## data:  PERCENT by GENDER
## t = 5.029, df = 70, p-value = 3.649e-06
## alternative hypothesis: true difference in means between group Male and group Female is not equal to 0
## 95 percent confidence interval:
##  3.295309 7.626913
## sample estimates:
##   mean in group Male mean in group Female 
##             18.84167             13.38056
t1$p.value
## [1] 3.648713e-06
t1$conf.int
## [1] 3.295309 7.626913
## attr(,"conf.level")
## [1] 0.95

Discussion

(For Two Sample t-test) Our decision should be to reject H0: Mu1 = Mu2 as the p < .05 and the 95% CI of the estimated population difference [3.295309 7.626913], which did not capture H0: Mu1 - Mu2 = 0. The results of the two-sample t-test were therefore statistically significant. This meant that the mean % of Male Smokers was significantly different from the Female Smokers.

References

The Dataset is taken From https://catalog.data.gov/dataset/adult-cigarette-and-tobacco-use-prevalence-31e2f Apart From that,Various Function and Ideas used in this assessment is taken by the help From Module Notes of Applied Analytics in canvas.