library(gapminder)
library(ggplot2)
data("gapminder")
income = gapminder$gdpPercap
life_exp = gapminder$lifeExp
graph = ggplot(data = gapminder, aes(x = income, y = life_exp))
graph = graph + geom_point(alpha = .25, color = "blue")
graph = graph + geom_line(color = "red", lwd = 1, stat = "smooth",
method = "loess")
graph = graph + xlab("Income") + ylab("Life Expectancy") +
ggtitle("Income vs Life Expectancy")
graph = graph + scale_y_continuous(trans = "identity", breaks = seq(0, 100, 10))
graph = graph + scale_x_continuous(trans = "identity",
breaks = seq(0, 50000, 10000),
limits = c(0, 50000))
graph
## `geom_smooth()` using formula = 'y ~ x'
## Warning: Removed 6 rows containing non-finite outside the scale range
## (`stat_smooth()`).
## Warning: Removed 6 rows containing missing values or values outside the scale range
## (`geom_point()`).
This graph shows that there is a trend where as the income increases,
the life expectancy can be expected to increase as well. I changed the
scale to 50000 to show the graph trend in more detailed. After 50000,
there is a couple of outliers that skews the graph to the right.
What is the type of the following vectors? Explain why they have that type.
a = c(1, NA + 1L, "C")
typeof(a)
## [1] "character"
str(a)
## chr [1:3] "1" NA "C"
The vector type is characters and this is due to R converting the numerical values to character type. There are three objects in the vector and each of the objects have a character type.
b = c(1L / 0, NA)
typeof(b)
## [1] "double"
str(b)
## num [1:2] Inf NA
This vector has a double type which is the default type for numerical values in R. The first value in the vector is positive infinity and the second value is NA. Since you are able to do arithmetic with NA it will be the same character type as infinity, meaning that both values will have the default double type. This can be seen by running the code typeof(NA + 1) which will give double even though the output is NA.
c = c(1:3, 5)
typeof(c)
## [1] "double"
str(c)
## num [1:4] 1 2 3 5
This is a vector that contains 4 numerical objects that are each are of the type double. Again, the default data type for numerical objects is double and you have to change it to an integer (whole numbers) by either using the as.integer function or by inputting a L after a number such as 1L. This would indicate to R that 1 is an integer.
d = c(3L, NaN + 1L)
typeof(d)
## [1] "double"
str(d)
## num [1:2] 3 NaN
This vector contain two objects that are both of the object type double. The first object is the number 3 and the second object is not a number + 1. In order for R to add the two values together, it converts the 1L into a double so it can add it to the NaN. Then the vector is converted into a double overall.
e = c(NA, TRUE)
typeof(e)
## [1] "logical"
str(e)
## logi [1:2] NA TRUE
This vector has two different objects that are of the data type logical. R converts NA to a logical type, but the NA is not change to a true or false type.
forecast = c("partly cloudy", "sunny", "sunny", "partly cloudy", "partly cloudy"
, "sunny", "partly cloudy", "partly cloudy", "partly cloudy"
, "partly cloudy", "partly cloudy", "cloudy", "rain", "cloudy")
days = c(1:14)
weather = data.frame(days, forecast, stringsAsFactors = TRUE)
weather
str(weather)
## 'data.frame': 14 obs. of 2 variables:
## $ days : int 1 2 3 4 5 6 7 8 9 10 ...
## $ forecast: Factor w/ 4 levels "cloudy","partly cloudy",..: 2 4 4 2 2 4 2 2 2 2 ...
The data frame creates a list of equal lengths and each vector will contain the data. By setting strings as factors to true, it converted the character vector into factors. Since the weather forecast does not contain snow, I did not included it in the data, which is why there is only 4 levels instead of 5.