# Homework 2 - R Code
# Name: Christopher Benavidez
# UTSA ID: 01885448
# Due: October 12, 2026

# Setup and data
heartbpchol <- read.csv("C://Users/12108/OneDrive/Documents/csv file/heartbpchol.csv", header=TRUE)

#Problem 1. Two methods for measuring the level of vitamin B12 in red blood cells were compared. Blood sample were taken from 10 healthy adults, and, for each blood sample, the B12 level was determined using both methods. The resultant data are given below.

#At the 1% significance level, we are interested in testing the hypothesis that there is no difference between the two methods with regard to the measurement of the B12 level in red blood cells.

# (a) Use the Student's t-test to test the hypothesis above and check the normality assumption
method1 <- c(204, 238, 209, 277, 197, 226, 203, 131, 282, 76) # Enter the B12 measurements for method 1

method2 <- c(199, 230, 198, 253, 180, 209, 213, 167, 250, 82) # Enter the B12 measurements for method 2

difference <- method1 - method2 # Calculate the differences between the two methods for each adult

shapiro.test(difference) # Check normality assumption for the differences
## 
##  Shapiro-Wilk normality test
## 
## data:  difference
## W = 0.93646, p-value = 0.5144
t.test(method1, method2, paired = TRUE) # Perform a paired t-test to compare the two B12 measurement methods
## 
##  Paired t-test
## 
## data:  method1 and method2
## t = 1.0035, df = 9, p-value = 0.3418
## alternative hypothesis: true mean difference is not equal to 0
## 95 percent confidence interval:
##  -7.776641 20.176641
## sample estimates:
## mean difference 
##             6.2
# (b) Use the sign test to test the hypothesis above.
difference <- method1 - method2 # Calculate the differences between method 1 and method 2

difference # Display the differences for each adult
##  [1]   5   8  11  24  17  17 -10 -36  32  -6
sum(difference > 0) # Count the positive differences
## [1] 7
sum(difference < 0) # Count the negative differences
## [1] 3
binom.test(sum(difference > 0), length(difference), p = 0.5) # Perform an exact binomal test for the sign test
## 
##  Exact binomial test
## 
## data:  sum(difference > 0) and length(difference)
## number of successes = 7, number of trials = 10, p-value = 0.3438
## alternative hypothesis: true probability of success is not equal to 0.5
## 95 percent confidence interval:
##  0.3475471 0.9332605
## sample estimates:
## probability of success 
##                    0.7
# (c) Use the Wilcoxon rank sum test to test the hypothesis above.
wilcox.test(method1, method2) # Perform the Wilcoxon rank sum test comparing the two methods
## 
##  Wilcoxon rank sum exact test
## 
## data:  method1 and method2
## W = 54.5, p-value = 0.754
## alternative hypothesis: true location shift is not equal to 0
# (d) Do the three tests mentioned above give the same conclusion for this problem? If not, which one should be preferred in practice?

# All three tests five the same conclusion: we fail to reject the null hypothesis. Therefore, there is not a sufficient evidence at the 1% significance level that the two B12 measurement methods differ.


#Problem 2. The data below show the marital status and happiness of individuals who participated in the General Social Survey in 2006.
#This table is referred to as a contingency table, or a two-way table. Since it related two categories of data. The row variable is happiness, the column variable is marital status, and each box inside the table is referred to as a cell. We are interested in determining whether marital status are happiness associated.

# (a) Set up the null and alternative hypotheses.

# Null Hypothesis: Marital status and happiness are independent; there is no association between marital status and happiness.
# Alternative Hypothesis: Marital status and happiness are not independent; there is an association between marital status and happiness

# (b) Perform a Chi-square test for a contingency table in determining whether marital status are happiness associated at 5% significance level.

#Enter the happiness and marital status count
happiness <- matrix(c(600, 63, 112, 144, 720, 142, 355, 459, 93, 51, 119, 127), nrow = 3, byrow = TRUE)

#Add row names for happiness categories
rownames(happiness) <- c("Very Happy", "Pretty Happy", "Not Too Happy")

#Add column names for marital status categories
colnames(happiness) <- c("Married", "Widowed", "Divorced/Separated", "Never Married")

#Display the contingency table
happiness
##               Married Widowed Divorced/Separated Never Married
## Very Happy        600      63                112           144
## Pretty Happy      720     142                355           459
## Not Too Happy      93      51                119           127
#Perform the Chi-square test of independence
chisq.test(happiness)
## 
##  Pearson's Chi-squared test
## 
## data:  happiness
## X-squared = 224.12, df = 6, p-value < 2.2e-16
# (c) Clearly report your results for this problem.

# A Pearson Chi-square test of independence was conducted to determine whether marital status and happiness are associated. The test result was X-squared = 224.12, df = 6, p-value < 2.2e-16. At the 5% significance level, the p-value is less than 0.05, so we reject the null hypothesis. There is sufficient evidence to conclude that marital status and happiness are associated.


#Problem 3. The heartbpchol.csv data set contains continuous cholesterol (Cholesterol) and blood pressure status (BP_Status)(categrory: High/ Normal/ Optimal) for alive patients.
#For the heartbpchol data set, consider a one-way ANOVA model to identify differences between group cholesterol means. The normality assumption is reasonable, so you can proceed without testing normality.

# (a) Perform a one-way ANOVA for Cholesterol with BP_Status as the categorical predictor. Comment on statistical significance of BP_Status, the amount of variation described by the model, and whether or not the equal variance assumption can be trusted. 
anova.chol <- aov(Cholesterol ~ BP_Status, data = heartbpchol)

#Display the ANOVA results
summary(anova.chol)
##              Df  Sum Sq Mean Sq F value  Pr(>F)   
## BP_Status     2   25211   12605   6.671 0.00137 **
## Residuals   538 1016631    1890                   
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#Fit a linear model to calculate the R-squared value
lm.chol <- lm(Cholesterol ~ BP_Status, data = heartbpchol)

#Display the model summary
summary(lm.chol)
## 
## Call:
## lm(formula = Cholesterol ~ BP_Status, data = heartbpchol)
## 
## Residuals:
##      Min       1Q   Median       3Q      Max 
## -112.029  -31.029   -4.029   21.428  177.428 
## 
## Coefficients:
##                  Estimate Std. Error t value Pr(>|t|)    
## (Intercept)       240.572      2.873  83.748  < 2e-16 ***
## BP_StatusNormal   -11.543      3.996  -2.889  0.00402 ** 
## BP_StatusOptimal  -18.647      6.038  -3.088  0.00212 ** 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 43.47 on 538 degrees of freedom
## Multiple R-squared:  0.0242, Adjusted R-squared:  0.02057 
## F-statistic: 6.671 on 2 and 538 DF,  p-value: 0.001375
#Test whether the variances are equal across BP Status groups
bartlett.test(Cholesterol ~ BP_Status, data = heartbpchol)
## 
##  Bartlett test of homogeneity of variances
## 
## data:  Cholesterol by BP_Status
## Bartlett's K-squared = 1.3845, df = 2, p-value = 0.5004
#Perform Tukey's post-hoc test to compare all BP Status groups
TukeyHSD(anova.chol)
##   Tukey multiple comparisons of means
##     95% family-wise confidence level
## 
## Fit: aov(formula = Cholesterol ~ BP_Status, data = heartbpchol)
## 
## $BP_Status
##                      diff       lwr       upr     p adj
## Normal-High    -11.543481 -20.93394 -2.153023 0.0111929
## Optimal-High   -18.646679 -32.83690 -4.456460 0.0059898
## Optimal-Normal  -7.103198 -21.18815  6.981749 0.4624869
# (b) Comment on any significantly cholesterol means as determined by the post-hoc test comparing all pairwise differences. Specifically explain what that tells us about differences in cholesterol levels across blood pressure status groups, like which group has the highest or lowest mean values of Cholesterol. 

# The Tukey post-hoc test showed that the High blood pressure status group had significantly higher cholesterol levels than both the Normal group (p = 0.0112) and the Optimal group (p = 0.0060). However, there was no significant difference between the Optimal and Normal groups (p = 0.4625). Therefore, the High BP Status group has the highest mean cholesterol, while the Optimal and Normal groups have cholesterol means that are not significantly different from each other. 

#Problem 4. For this problem use the bupa.csv data set. The mean corpuscular volume and alkaline phosphase are blood tests thought to be sensitive to liver disorder related to excessive alcohol consumption. We assume that normality and independence assumptions are valid.
bupa.data <- read.csv("C://Users/12108/OneDrive/Documents/csv file/bupa.data", header=FALSE)

# (a) Perform a one-way ANOVA for mcv as a function of drinkgroup. Comment on significance of the drinkgroup, the amount of variation described by the model, and whether or not the equal variance assumption can be trusted.
#Add the variable names to the data
names(bupa.data) <- c("mcv", "alkphos", "sgpt", "sgot", "gammagt", "drinkgroup", "selector")

#Create drink group from the number of drinks
bupa.data$drinkgroup <- cut(bupa.data$drinkgroup,
                            breaks = c(-Inf, 0.5, 2, 5, 8, Inf),
                            labels = c(1, 2, 3, 4, 5))

#Check the number of observations in each drink group
table(bupa.data$drinkgroup)
## 
##   1   2   3   4   5 
## 117  52  88  67  21
#Perform a one-way ANOVA for mcv by drinkgroup
anova.mcv <- aov(mcv ~ drinkgroup, data = bupa.data)

#Display the ANOVA results
summary(anova.mcv)
##              Df Sum Sq Mean Sq F value   Pr(>F)    
## drinkgroup    4    733  183.29   10.26 7.43e-08 ***
## Residuals   340   6073   17.86                     
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#Fit a linear model for MCV by drink group
lm.mcv <- lm(mcv ~ drinkgroup, data = bupa.data)

#Display the model summary
summary(lm.mcv)
## 
## Call:
## lm(formula = mcv ~ drinkgroup, data = bupa.data)
## 
## Residuals:
##      Min       1Q   Median       3Q      Max 
## -23.7778  -2.7159  -0.0192   2.4762  12.9808 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept)  88.7778     0.3907 227.213  < 2e-16 ***
## drinkgroup2   1.2415     0.7044   1.762 0.078891 .  
## drinkgroup3   0.9381     0.5964   1.573 0.116625    
## drinkgroup4   3.7446     0.6475   5.783 1.66e-08 ***
## drinkgroup5   3.7460     1.0016   3.740 0.000216 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 4.226 on 340 degrees of freedom
## Multiple R-squared:  0.1077, Adjusted R-squared:  0.09722 
## F-statistic: 10.26 on 4 and 340 DF,  p-value: 7.429e-08
#Test whether the MCV variances are equal across drink groups
bartlett.test(mcv ~ drinkgroup, data = bupa.data)
## 
##  Bartlett test of homogeneity of variances
## 
## data:  mcv by drinkgroup
## Bartlett's K-squared = 2.8248, df = 4, p-value = 0.5876
# (b) Perform a one-way ANOVA for alkphos as a function of drinkgroup. Comment on statistical significance of the drinkgroup, the amount of variation described by the model, and whether or not the equal variance assumption can be trusted.

#Perform a one-way ANOVA for alkphos by drink group
anova.alkphos <- aov(alkphos ~ drinkgroup, data = bupa.data)

#Display the ANOVA results
summary(anova.alkphos)
##              Df Sum Sq Mean Sq F value  Pr(>F)   
## drinkgroup    4   4946  1236.4   3.792 0.00495 **
## Residuals   340 110858   326.1                   
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#Fit a linear model for alkphos by drink group
lm.alkphos <- lm(alkphos ~ drinkgroup, data = bupa.data)

#Display the model summary
summary(lm.alkphos)
## 
## Call:
## lm(formula = alkphos ~ drinkgroup, data = bupa.data)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -47.761 -12.612  -2.761   9.388  67.295 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept)   70.761      1.669  42.388  < 2e-16 ***
## drinkgroup2   -2.645      3.009  -0.879  0.38003    
## drinkgroup3   -4.056      2.548  -1.592  0.11233    
## drinkgroup4   -1.149      2.766  -0.415  0.67823    
## drinkgroup5   12.573      4.279   2.938  0.00353 ** 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 18.06 on 340 degrees of freedom
## Multiple R-squared:  0.04271,    Adjusted R-squared:  0.03144 
## F-statistic: 3.792 on 4 and 340 DF,  p-value: 0.00495
#Test whether the alkphos variance are equal across drink groups
bartlett.test(alkphos ~ drinkgroup, data = bupa.data)
## 
##  Bartlett test of homogeneity of variances
## 
## data:  alkphos by drinkgroup
## Bartlett's K-squared = 5.4205, df = 4, p-value = 0.2468
# (c) Perform post-hoc tests for models in a) and b). Comment on any similarities or differences you observe from their results.

#Perform Tukey's post-hoc test for MCV
TukeyHSD(anova.mcv)
##   Tukey multiple comparisons of means
##     95% family-wise confidence level
## 
## Fit: aov(formula = mcv ~ drinkgroup, data = bupa.data)
## 
## $drinkgroup
##             diff          lwr      upr     p adj
## 2-1  1.241452991 -0.690316778 3.173223 0.3973587
## 3-1  0.938131313 -0.697363908 2.573627 0.5157202
## 4-1  3.744610282  1.968846495 5.520374 0.0000002
## 5-1  3.746031746  0.999127120 6.492936 0.0020039
## 3-2 -0.303321678 -2.330666517 1.724023 0.9940309
## 4-2  2.503157290  0.361050967 4.645264 0.0127884
## 5-2  2.504578755 -0.492214115 5.501372 0.1499436
## 4-3  2.806478969  0.927189295 4.685769 0.0005031
## 5-3  2.807900433 -0.007037875 5.622839 0.0509321
## 5-4  0.001421464 -2.897262735 2.900106 1.0000000
#Perform Tukey's post-hoc test for alkphos
TukeyHSD(anova.alkphos)
##   Tukey multiple comparisons of means
##     95% family-wise confidence level
## 
## Fit: aov(formula = alkphos ~ drinkgroup, data = bupa.data)
## 
## $drinkgroup
##          diff         lwr       upr     p adj
## 2-1 -2.645299 -10.8987258  5.608127 0.9045115
## 3-1 -4.056138 -11.0437411  2.931464 0.5035761
## 4-1 -1.148743  -8.7356393  6.438152 0.9937509
## 5-1 12.572650   0.8365845 24.308715 0.0288892
## 3-2 -1.410839 -10.0726073  7.250929 0.9917361
## 4-2  1.496556  -7.6555274 10.648639 0.9916117
## 5-2 15.217949   2.4142438 28.021654 0.0107154
## 4-3  2.907395  -5.1218122 10.936602 0.8583931
## 5-3 16.628788   4.6020510 28.655525 0.0016498
## 5-4 13.721393   1.3368544 26.105932 0.0214375
#The post-hoc results showed that the significant differences between drink groups were different fot MCV and alkphos. For MCV, drink group 4 had significantly higher MCV than groups 1,2, and 3, and group 5 had significantly higher MCV than group 1. For alkphos, group 5 had significantly higher alkphos than group 1, 3, and 4. Therefore, although both variables showed significant differences among drink groups in the overall ANOVA, the specific drink groups that differed were not the same.