Sukhvir Singh Chahal (s4019923) and Ruhan Ahmed Malik (s4024339)
Last updated: 15th october 2023
The dataset offers a thorough perspective on smoking prevalence in the United States Covering the years from 1984 to 2019 Our analysis centers on two important variables: “GENDER,” which categorizes the data into “Male,” and “Female,” and “PERCENT” reflecting smoking percentages within these gender groups.
The main objective is to gain insights into the patterns of smoking behavior within the U.S. population and to explore whether there are significant difference in smoking rates between Male and Female Our analysis is centered on two separate t-tests: a two-sample t-test to compare male and female smoking percentages.
Two-Sample t-test for Male and Female Smoking Percentages: a two-sample t-test, we would try to determine whether there exists a statistically significant smoking behavior between these two gender categories.
The Dataset is taken From https://catalog.data.gov/dataset/adult-cigarette-and-tobacco-use-prevalence-31e2f The dataset provides a comprehensive perspective on smoking prevalence and associated confidence intervals, covering the years from 1984 to 2019 in the United states. Within this dataset, we focus our analysis on two key variables:
GENDER: This variable classifies the data into three distinct groups, namely “Male,” and “Female”.
PERCENT: It represents the percentage of individuals within specific gender groups who have reported their smoking behavior.
adult_smoking <- read.csv("adult-smoking-prevalence.csv")
adult_smoking$GENDER<- adult_smoking$GENDER %>% factor(levels=c("Male","Female","Total"))
Male_smoking <- adult_smoking %>% filter(GENDER =="Male")
Female_smoking <- adult_smoking %>% filter(GENDER =="Female")
Adult_smoking_New <- union(Male_smoking,Female_smoking)
str(Adult_smoking_New)## 'data.frame': 72 obs. of 6 variables:
## $ YEAR : int 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 ...
## $ COMPARISON: chr "Definition Expanded in 1996" "Definition Expanded in 1996" "Definition Expanded in 1996" "Definition Expanded in 1996" ...
## $ GENDER : Factor w/ 3 levels "Male","Female",..: 1 1 1 1 1 1 1 1 1 1 ...
## $ PERCENT : num 26.9 28.2 28 22.8 25.6 22.5 21.1 22.7 22.3 20.7 ...
## $ LOWER95 : num 22.8 24.6 24.6 19.8 22.9 20 18.8 20.5 20.4 19.3 ...
## $ UPPER95 : num 31 31.8 31.4 25.8 28.2 25 23.3 25 24.2 22.1 ...
## Rows: 36
## Columns: 6
## $ YEAR <int> 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993,…
## $ COMPARISON <chr> "Definition Expanded in 1996", "Definition Expanded in 1996…
## $ GENDER <fct> Male, Male, Male, Male, Male, Male, Male, Male, Male, Male,…
## $ PERCENT <dbl> 26.9, 28.2, 28.0, 22.8, 25.6, 22.5, 21.1, 22.7, 22.3, 20.7,…
## $ LOWER95 <dbl> 22.8, 24.6, 24.6, 19.8, 22.9, 20.0, 18.8, 20.5, 20.4, 19.3,…
## $ UPPER95 <dbl> 31.0, 31.8, 31.4, 25.8, 28.2, 25.0, 23.3, 25.0, 24.2, 22.1,…
## Rows: 36
## Columns: 6
## $ YEAR <int> 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993,…
## $ COMPARISON <chr> "Definition Expanded in 1996", "Definition Expanded in 1996…
## $ GENDER <fct> Female, Female, Female, Female, Female, Female, Female, Fem…
## $ PERCENT <dbl> 23.0, 25.2, 23.3, 20.1, 19.9, 19.8, 17.9, 15.8, 17.8, 15.8,…
## $ LOWER95 <dbl> 19.7, 22.1, 20.5, 17.7, 17.8, 17.7, 15.9, 14.0, 16.2, 14.6,…
## $ UPPER95 <dbl> 26.3, 28.2, 26.0, 22.6, 22.0, 22.0, 19.8, 17.6, 19.4, 16.9,…
## Rows: 72
## Columns: 6
## $ YEAR <int> 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993,…
## $ COMPARISON <chr> "Definition Expanded in 1996", "Definition Expanded in 1996…
## $ GENDER <fct> Male, Male, Male, Male, Male, Male, Male, Male, Male, Male,…
## $ PERCENT <dbl> 26.9, 28.2, 28.0, 22.8, 25.6, 22.5, 21.1, 22.7, 22.3, 20.7,…
## $ LOWER95 <dbl> 22.8, 24.6, 24.6, 19.8, 22.9, 20.0, 18.8, 20.5, 20.4, 19.3,…
## $ UPPER95 <dbl> 31.0, 31.8, 31.4, 25.8, 28.2, 25.0, 23.3, 25.0, 24.2, 22.1,…
Male_smoking %>% summarise(Min = min(PERCENT),
Q1 = quantile(PERCENT,probs = .25),
Median = median(PERCENT),
Q3 = quantile(PERCENT,probs = .75),
Max = max(PERCENT),
Mean = mean(PERCENT),
SD = sd(PERCENT))Female_smoking %>% summarise(Min = min(PERCENT),
Q1 = quantile(PERCENT,probs = .25),
Median = median(PERCENT),
Q3 = quantile(PERCENT,probs = .75),
Max = max(PERCENT),
Mean = mean(PERCENT),
SD = sd(PERCENT)) (For Two Sample t-test) Through Descriptive Statistics, The mean percent of Male smokers is Not equal to that of Female Group, but we are not sure yet whether it is significantly different or not.
## [1] 2 3
## [1] 2 3
According to qq plot and box plots, For both groups, Percent smokers are in normal distribution and moreover Both groups have sample size greater than 30, so the distribution in both groups are normal. As p>0.05, population variances are homogeneous(by levene Test) Thus sampling distribution will approximate a normal distribution and we can apply the two-sample t-test.
(For Two Sample t-test)
\[H_0: \mu_1 = \mu_2\] \[H_A: \mu_1 \ne \mu_2\]
alpha <- 0.05
Average_Smoker_claim <- 11.5
t1 <- t.test(PERCENT ~ GENDER,
data = Adult_smoking_New,
var.equal = TRUE,
alternative = "two.sided")
t1##
## Two Sample t-test
##
## data: PERCENT by GENDER
## t = 5.029, df = 70, p-value = 3.649e-06
## alternative hypothesis: true difference in means between group Male and group Female is not equal to 0
## 95 percent confidence interval:
## 3.295309 7.626913
## sample estimates:
## mean in group Male mean in group Female
## 18.84167 13.38056
## [1] 3.648713e-06
## [1] 3.295309 7.626913
## attr(,"conf.level")
## [1] 0.95
(For Two Sample t-test) Our decision should be to reject H0: Mu1 = Mu2 as the p < .05 and the 95% CI of the estimated population difference [3.295309 7.626913], which did not capture H0: Mu1 - Mu2 = 0. The results of the two-sample t-test were therefore statistically significant. This meant that the mean % of Male Smokers was significantly different from the Female Smokers.
The Dataset is taken From https://catalog.data.gov/dataset/adult-cigarette-and-tobacco-use-prevalence-31e2f Apart From that,Various Function and Ideas used in this assessment is taken by the help From Module Notes of Applied Analytics in canvas.