flint = read.csv("~/Downloads/flint-2.csv")
sum(flint$Pb >= 15)/nrow(flint)
## [1] 0.04436229
The proportion of dangerous lead levels was about 0.044.
northRegion = flint$Region == "North"
mean(flint$Cu[flint$Region == "North"]) # Mean copper level for testing sites in the North Region
## [1] 44.6424
The mean copper level for test sites in the North region is 44.6424.
mean(flint[flint$Pb >= 15, 'Cu'])
## [1] 305.8333
The mean copper level for test sites only with dangerous levels is 305.8333.
mean(flint$Pb) # Mean lead level
## [1] 3.383272
mean(flint$Cu) # Mean copper level
## [1] 54.58102
The mean lead level is about 3.383, and the mean copper level is about 54.581.
boxplot(flint$Pb, main = "Flint Lead Levels")
The mean does not seem like a good measure of center because there appears to be a lot of outliers on the boxplot. It would be better to use the median to avoid this skew.
life <- read.table("http://www.stat.ucla.edu/~nchristo/statistics12/countries_life.txt", header = TRUE)
plot(x = life$Income, y = life$Life)
A higher life expectancy seems to be correlated with a higher income to a certain extent. Though, from $1000 on, the life expectancy plateaus and seems fairly constant despite increasing income.
library(mosaic)
## Registered S3 method overwritten by 'mosaic':
## method from
## fortify.SpatialPolygonsDataFrame ggplot2
##
## The 'mosaic' package masks several functions from core packages in order to add
## additional features. The original behavior of these functions should not be affected by this.
##
## Attaching package: 'mosaic'
## The following objects are masked from 'package:dplyr':
##
## count, do, tally
## The following object is masked from 'package:Matrix':
##
## mean
## The following object is masked from 'package:ggplot2':
##
## stat
## The following objects are masked from 'package:stats':
##
## binom.test, cor, cor.test, cov, fivenum, IQR, median, prop.test,
## quantile, sd, t.test, var
## The following objects are masked from 'package:base':
##
## max, mean, min, prod, range, sample, sum
boxplot(life$Income, main = "Boxplot of the Income")
histogram(life$Income, main = "Histogram of the Income")
There seem to be outliers past 300.
higherIncome = life[life$Income >= 1000,]
head(higherIncome)
## Country Life Income
## 1 AUSTRALIA 71.0 3426
## 2 AUSTRIA 70.4 3350
## 3 BELGIUM 70.6 3346
## 4 CANADA 72.0 4751
## 5 DENMARK 73.3 5029
## 6 FINLAND 69.8 3312
lowerIncome = life[life$Income < 1000,]
head(lowerIncome)
## Country Life Income
## 15 PORTUGAL 68.1 956
## 20 ALGERIA 50.7 430
## 21 ECUADOR 52.3 360
## 22 INDONESIA 47.5 110
## 24 IRAQ 51.6 560
## 26 NIGERIA 36.9 180
plot(x = lowerIncome$Income, y = lowerIncome$Life)
cor(lowerIncome$Income, lowerIncome$Life)
## [1] 0.752886
maas = read.table("http://www.stat.ucla.edu/~nchristo/statistics12/soil.txt", header = TRUE)
summary(maas$lead)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 37.0 72.5 123.0 153.4 207.0 654.0
summary(maas$zinc)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 113.0 198.0 326.0 469.7 674.5 1839.0
histogram(maas$lead)
histogram(log(maas$lead))
plot(x = log(maas$lead), y = log(maas$zinc))
cor(log(maas$lead), log(maas$zinc))
## [1] 0.9671621
The log of these values is very highly correlated, with a strong, positive association.
lead_colors = c("green", "yellow", "orange","red")
lead_levels = cut(maas$lead, c(0.150,400,655))
as.numeric(lead_levels)
## [1] 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
## [38] 1 1 2 1 1 1 1 1 1 1 1 1 1 1 1 2 2 2 1 1 1 2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
## [75] 1 1 1 1 2 2 1 2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
## [112] 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
## [149] 1 1 1 1 1 1 1
plot(maas$x, maas$y, xlab="Longitude",
ylab="Latitude", main="Lead Levels in Maas River", "n")
points(maas$x, maas$y, cex=maas$lead/mean(maas$lead), col=lead_colors[as.numeric(lead_levels)], pch=19)
LA <- read.table("http://www.stat.ucla.edu/~nchristo/statistics12/la_data.txt", header = TRUE)
library(maps)
plot(x=LA$Longitude, y=LA$Latitude, xlab = "Longitude", ylab = "Latitude", main = "City of LA Neighborhood Centers")
map("county", "California", add = TRUE)
subsettedLA = LA[LA$Schools != 0,]
plot(x=subsettedLA$Income, y = subsettedLA$Schools, xlab = "Income", ylab = "Schools", main = "School Performance versus Income")
Income and School performance seem to be closely related up to an income of 100,000, and then more loosely correlated.