mydata <- read.csv("train.csv")
library(psych)
describe(mydata)
## vars n mean sd median trimmed mad min max range
## PassengerId 1 891 446.00 257.35 446.00 446.00 330.62 1.00 891.00 890.00
## Survived 2 891 0.38 0.49 0.00 0.35 0.00 0.00 1.00 1.00
## Pclass 3 891 2.31 0.84 3.00 2.39 0.00 1.00 3.00 2.00
## Name* 4 891 446.00 257.35 446.00 446.00 330.62 1.00 891.00 890.00
## Sex* 5 891 1.65 0.48 2.00 1.68 0.00 1.00 2.00 1.00
## Age 6 714 29.70 14.53 28.00 29.27 13.34 0.42 80.00 79.58
## SibSp 7 891 0.52 1.10 0.00 0.27 0.00 0.00 8.00 8.00
## Parch 8 891 0.38 0.81 0.00 0.18 0.00 0.00 6.00 6.00
## Ticket* 9 891 339.52 200.83 338.00 339.65 268.35 1.00 681.00 680.00
## Fare 10 891 32.20 49.69 14.45 21.38 10.24 0.00 512.33 512.33
## Cabin* 11 891 18.63 38.14 1.00 8.29 0.00 1.00 148.00 147.00
## Embarked* 12 891 3.53 0.80 4.00 3.66 0.00 1.00 4.00 3.00
## skew kurtosis se
## PassengerId 0.00 -1.20 8.62
## Survived 0.48 -1.77 0.02
## Pclass -0.63 -1.28 0.03
## Name* 0.00 -1.20 8.62
## Sex* -0.62 -1.62 0.02
## Age 0.39 0.16 0.54
## SibSp 3.68 17.73 0.04
## Parch 2.74 9.69 0.03
## Ticket* 0.00 -1.28 6.73
## Fare 4.77 33.12 1.66
## Cabin* 2.09 3.07 1.28
## Embarked* -1.27 -0.16 0.03
str(mydata)
## 'data.frame': 891 obs. of 12 variables:
## $ PassengerId: int 1 2 3 4 5 6 7 8 9 10 ...
## $ Survived : int 0 1 1 1 0 0 0 0 1 1 ...
## $ Pclass : int 3 1 3 1 3 3 1 3 3 2 ...
## $ Name : chr "Braund, Mr. Owen Harris" "Cumings, Mrs. John Bradley (Florence Briggs Thayer)" "Heikkinen, Miss. Laina" "Futrelle, Mrs. Jacques Heath (Lily May Peel)" ...
## $ Sex : chr "male" "female" "female" "female" ...
## $ Age : num 22 38 26 35 35 NA 54 2 27 14 ...
## $ SibSp : int 1 1 0 1 0 0 0 3 0 1 ...
## $ Parch : int 0 0 0 0 0 0 0 1 2 0 ...
## $ Ticket : chr "A/5 21171" "PC 17599" "STON/O2. 3101282" "113803" ...
## $ Fare : num 7.25 71.28 7.92 53.1 8.05 ...
## $ Cabin : chr "" "C85" "" "C123" ...
## $ Embarked : chr "S" "C" "S" "S" ...
Q1. a. What are the types of variable (quantitative / qualitative) and levels of measurement (nominal / ordinal / interval / ratio) for PassengerId and Age?
The type of PassengerId and Age is Quantative. The level of measurment for the PassengerId variable is nominal, and for the Age variable is ratio.
Q1b. Which variable has the most missing observations? You could have Googled for “Count NA values in R for all columns in a dataframe” or something like that.
colSums(is.na(mydata))
## PassengerId Survived Pclass Name Sex Age
## 0 0 0 0 0 177
## SibSp Parch Ticket Fare Cabin Embarked
## 0 0 0 0 0 0
The age variable is the only variable with missing observations. The number of missing observations is 177.
Q2. Impute missing observations for Age, SibSp, and Parch with the column median (ordinal, interval, or ratio), or the column mode. To do so, use something like this: mydata\(Age[is.na(mydata\)Age)] <- median(mydata\(Age, na.rm=TRUE) . You can read up on indexing in dataframe (i.e. "dataframe_name\)variable_name”).
imputeAge <- mydata$Age[is.na(mydata$Age)] <- median(mydata$Age,na.rm=TRUE)
print(imputeAge)
## [1] 28
The median age is 28.
Q3. Install the psych package in R: install.packages(‘pscyh’). Invoke the package using library(psych). Then provide descriptive statistics for Age, SibSp, and Parch (e.g., describe(mydata$Age). Please comment on what you observe from the summary statistics.
describe(mydata$Age)
## vars n mean sd median trimmed mad min max range skew kurtosis se
## X1 1 891 29.36 13.02 28 28.83 8.9 0.42 80 79.58 0.51 0.97 0.44
describe(mydata$SibSp)
## vars n mean sd median trimmed mad min max range skew kurtosis se
## X1 1 891 0.52 1.1 0 0.27 0 0 8 8 3.68 17.73 0.04
describe(mydata$Parch)
## vars n mean sd median trimmed mad min max range skew kurtosis se
## X1 1 891 0.38 0.81 0 0.18 0 0 6 6 2.74 9.69 0.03
The age among the passengers vary from around 5 months old youngest to 80 being the oldest passenger. The median is 28 and the standard deviation is 13, meaning that approximately 68% of the passengers are between 15 to 41 years old. An interesting fact about the SibSp variable is the max is 8, meaning that some passenger, or passengers, has 8 siblings/spouses on the ship. Regarding the Parch variable, someone has 6 parents/children on the ship.
Q4. Provide a cross-tabulation of Survived and Sex (e.g., table(mydata\(Survived, mydata\)Sex). What do you notice?
table(mydata$Survived,mydata$Sex)
##
## female male
## 0 81 468
## 1 233 109
I notice that the survival proportion among women is much higher than the survival proportion among men. The survival proportion for women was 0.74 while for men it was 0.18. I wonder what are the causes for that. My guess is that many men tried to save their family, or other people, and died in the process. Perhaps the ship’s crew was mostly consisted in men and they were evacuating women first, and then their fellow males.
Q5. Provide notched boxplots for Survived and Age (e.g., boxplot(mydata\(Age~mydata\)Survived, notch=TRUE, horizontal=T). What do you notice?
boxplot(mydata$Age~mydata$Survived,notch=TRUE,horizontal=T,
names= c("Deceased","Survived"),
xlab="Passenger age in years",
ylab= "Survival Status"
)
I notice that the data for the people who survived is more varied, shown by the bigger IQR and and standard deviation. There are less outliers, with one extremely old surviving passenger, aged more than 80 that is far from the rest of the survivors’ ages. The data for the deceased passengers is less varied, suggesting they were in similar ages. The IQR and standard deviation is smaller, and there are many outliers, appearing from both sides of box.