Exercise 1

Part a

flint = read.csv("~/Downloads/flint-2.csv")

Part b

sum(flint$Pb >= 15)/nrow(flint)
## [1] 0.04436229

The proportion of dangerous lead levels was about 0.044.

Part c

northRegion = flint$Region == "North"
mean(flint$Cu[flint$Region == "North"]) # Mean copper level for testing sites in the North Region
## [1] 44.6424

The mean copper level for test sites in the North region is 44.6424.

Part d

mean(flint[flint$Pb >= 15, 'Cu']) 
## [1] 305.8333

The mean copper level for test sites only with dangerous levels is 305.8333.

Part e

mean(flint$Pb) # Mean lead level
## [1] 3.383272
mean(flint$Cu) # Mean copper level
## [1] 54.58102

The mean lead level is about 3.383, and the mean copper level is about 54.581.

Part f

boxplot(flint$Pb, main = "Flint Lead Levels")

Part g

The mean does not seem like a good measure of center because there appears to be a lot of outliers on the boxplot. It would be better to use the median to avoid this skew.

Exercise 2

life <- read.table("http://www.stat.ucla.edu/~nchristo/statistics12/countries_life.txt", header = TRUE)

Part a

plot(x = life$Income, y = life$Life)

A higher life expectancy seems to be correlated with a higher income to a certain extent. Though, from $1000 on, the life expectancy plateaus and seems fairly constant despite increasing income.

Part b

library(mosaic)
## Registered S3 method overwritten by 'mosaic':
##   method                           from   
##   fortify.SpatialPolygonsDataFrame ggplot2
## 
## The 'mosaic' package masks several functions from core packages in order to add 
## additional features.  The original behavior of these functions should not be affected by this.
## 
## Attaching package: 'mosaic'
## The following objects are masked from 'package:dplyr':
## 
##     count, do, tally
## The following object is masked from 'package:Matrix':
## 
##     mean
## The following object is masked from 'package:ggplot2':
## 
##     stat
## The following objects are masked from 'package:stats':
## 
##     binom.test, cor, cor.test, cov, fivenum, IQR, median, prop.test,
##     quantile, sd, t.test, var
## The following objects are masked from 'package:base':
## 
##     max, mean, min, prod, range, sample, sum
boxplot(life$Income, main = "Boxplot of the Income")

histogram(life$Income, main = "Histogram of the Income")

There seem to be outliers past 300.

Part c

higherIncome = life[life$Income >= 1000,]
head(higherIncome)
##     Country Life Income
## 1 AUSTRALIA 71.0   3426
## 2   AUSTRIA 70.4   3350
## 3   BELGIUM 70.6   3346
## 4    CANADA 72.0   4751
## 5   DENMARK 73.3   5029
## 6   FINLAND 69.8   3312
lowerIncome = life[life$Income < 1000,]
head(lowerIncome)
##      Country Life Income
## 15  PORTUGAL 68.1    956
## 20   ALGERIA 50.7    430
## 21   ECUADOR 52.3    360
## 22 INDONESIA 47.5    110
## 24      IRAQ 51.6    560
## 26   NIGERIA 36.9    180

Part d

plot(x = lowerIncome$Income, y = lowerIncome$Life)

cor(lowerIncome$Income, lowerIncome$Life)
## [1] 0.752886

Exercise 3

maas = read.table("http://www.stat.ucla.edu/~nchristo/statistics12/soil.txt", header = TRUE)

a

summary(maas$lead)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##    37.0    72.5   123.0   153.4   207.0   654.0
summary(maas$zinc)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   113.0   198.0   326.0   469.7   674.5  1839.0

b

histogram(maas$lead)

histogram(log(maas$lead))

Part c

plot(x = log(maas$lead), y = log(maas$zinc))

cor(log(maas$lead), log(maas$zinc))
## [1] 0.9671621

The log of these values is very highly correlated, with a strong, positive association.

Part d

lead_colors = c("green", "yellow", "orange","red")

lead_levels = cut(maas$lead, c(0.150,400,655))
                  
as.numeric(lead_levels)
##   [1] 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
##  [38] 1 1 2 1 1 1 1 1 1 1 1 1 1 1 1 2 2 2 1 1 1 2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
##  [75] 1 1 1 1 2 2 1 2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
## [112] 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
## [149] 1 1 1 1 1 1 1
plot(maas$x, maas$y, xlab="Longitude",
ylab="Latitude", main="Lead Levels in Maas River", "n")

points(maas$x, maas$y, cex=maas$lead/mean(maas$lead), col=lead_colors[as.numeric(lead_levels)], pch=19)

Exercise 4

LA <- read.table("http://www.stat.ucla.edu/~nchristo/statistics12/la_data.txt", header = TRUE)

Part a

library(maps)
plot(x=LA$Longitude, y=LA$Latitude, xlab = "Longitude", ylab = "Latitude", main = "City of LA Neighborhood Centers")

map("county", "California", add = TRUE)

Part b

subsettedLA = LA[LA$Schools != 0,]

plot(x=subsettedLA$Income, y = subsettedLA$Schools, xlab = "Income", ylab = "Schools", main = "School Performance versus Income")

Income and School performance seem to be closely related up to an income of 100,000, and then more loosely correlated.