Airbnb<-read.csv("airbnb_134.csv", header=TRUE, sep=",") 
library(latticeExtra) 
## Loading required package: lattice
summary(Airbnb$bathrooms) 
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   1.000   1.000   1.000   1.420   1.875   4.500
summary(Airbnb$bedrooms) 
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##     0.0     1.0     1.0     1.5     2.0     6.0
  1. The two quantitative values we think may be related to price are bedrooms and bathrooms.

2a) The mean and median are both measures of center. The mean of the bathrooms is 1.420 and the median is 1.000. As for the bedrooms, the mean is 1.5 and the median is 1.0. We believe that both of the variables seems to be squewed right since the mean is greater than the median for both of them.

2b) For the bathroom variable, the standard deviation is 0.791279, the range is 3.5, and the interquartile range is 0.875.

For the bedroom variable, the standard deviation is 1.19949, the range is 6 and the interquartile range is 1.

Because the two variables have a skewed distribution, IQR is preferred for measure of spread over standard deviation which is more suitable for normal distribution.

bwplot(~Airbnb$bathrooms) 

bwplot(~Airbnb$bedrooms) 

histogram(~Airbnb$bathrooms) 

histogram(~Airbnb$bedrooms) 

3ai) Both of the data are skewed right because most of the data is gathers towards the left end of the x-axis with fewer data points toward the right.

3aii) There are 2 outliers for the bathroom variable and 3 outliers for the bedroom variable. These can be identified in the box plot by the dots past the far end of the right whiskers. Outliers are data that do not fall with the rest of the data points and do not fit the general pattern.

bwplot(~Airbnb$price) 

histogram(~Airbnb$price) 

  1. The center is around 100-200 and the mean would be greater than the median. The majority of the data falls around the center. The spread of the large chunk of data is about 400 but if you take outliers into account that increases the spread to about 800. From the histogram, it looks symetrical besides the outliers in the 700-800 range.
xyplot(price~bedrooms, data=Airbnb, type=c("p","r")) 

xyplot(price~bathrooms, data=Airbnb, type=c("p","r")) 

5) There seems to be a positive and linear relationship with the bedroom and price variable as well as the bathroom and price variable, but it’s a weak association. We no do not detect a non-linear relationship between price and our variables. Therefore, there is not a trend to warrant the transformation of our data since it is linear.

reg = lm(price ~ bathrooms + bedrooms, data=Airbnb) 
summary(reg) 
## 
## Call:
## lm(formula = price ~ bathrooms + bedrooms, data = Airbnb)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -248.60  -40.04  -14.67   24.96  530.54 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)   
## (Intercept)   -59.12      36.37  -1.626  0.11074   
## bathrooms     125.29      43.58   2.875  0.00605 **
## bedrooms       26.50      28.75   0.922  0.36135   
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 118.7 on 47 degrees of freedom
## Multiple R-squared:  0.5473, Adjusted R-squared:  0.528 
## F-statistic: 28.41 on 2 and 47 DF,  p-value: 8.156e-09

QUESTION 2 6) y = -59.12 + 125.29x + 26.50x2 7) The R^2 value is 0.5473. That means that half of the variance in the outcome if explained by the model. This is a moderately good relationship sine only ~50% of the data corresponds with the regression line.

plot(reg,2) 

histogram(reg$residuals) 

8b) No, the data looks roughly linear in the ggplot but it starts to curve once the outliers are present in the data. When looking at the histogram, it makes it clear that the data is linear but there is an outlier that may be throwing the ggplot off.

plot(reg,1) 

8d) There doesn’t seem to be a pattern in the residuals vs. fitted graph. It is not linear due to the curved nature of the graph and does not take on another other pattern.

  1. There are outliers on the far right in each of the variable datasets. No, there are not that many outliers that are affecting the data to be removed. Since there are multiple in the areas we can determine they were not typos and are truly supposed to be apart of the dataset. Additionally, if we think about it logically, having a four bedroom/bathroom Airbnb is not uncommon.
sqrtyourbathrooms = sqrt(Airbnb$bathrooms) 
sqrtyourbedrooms = sqrt(Airbnb$bedrooms) 
regg = lm(price ~ sqrtyourbathrooms + sqrtyourbedrooms, data=Airbnb) 
summary(regg)
## 
## Call:
## lm(formula = price ~ sqrtyourbathrooms + sqrtyourbedrooms, data = Airbnb)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -258.19  -38.86  -12.59   25.66  506.80 
## 
## Coefficients:
##                   Estimate Std. Error t value Pr(>|t|)    
## (Intercept)        -330.43      78.95  -4.185 0.000124 ***
## sqrtyourbathrooms   394.59      89.83   4.392 6.34e-05 ***
## sqrtyourbedrooms     28.70      46.87   0.612 0.543254    
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 124.9 on 47 degrees of freedom
## Multiple R-squared:  0.4981, Adjusted R-squared:  0.4768 
## F-statistic: 23.32 on 2 and 47 DF,  p-value: 9.208e-08
  1. The t value for bathrooms transformed is 4.392. Since this is a right tailed test we determined that the p value for bathrooms is 6.34e-05 which is less than 0.05. This means that the variable is significant in the model and is not due to random chance. The t value for bedrooms transformed is 0.612. Since this is a right tailed test we determined that the p value for bedrooms is 0.543254 is greater than 0.05. This means that the variable is not significant.

  2. The least significant variable is bedrooms and can be dropped.

reg1 = lm(price ~ sqrtyourbathrooms, data=Airbnb) 
summary(reg1)
## 
## Call:
## lm(formula = price ~ sqrtyourbathrooms, data = Airbnb)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -256.52  -45.92   -9.67   32.19  505.82 
## 
## Coefficients:
##                   Estimate Std. Error t value Pr(>|t|)    
## (Intercept)        -343.71      75.42  -4.557 3.58e-05 ***
## sqrtyourbathrooms   433.38      63.29   6.847 1.26e-08 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 124.1 on 48 degrees of freedom
## Multiple R-squared:  0.4941, Adjusted R-squared:  0.4836 
## F-statistic: 46.88 on 1 and 48 DF,  p-value: 1.262e-08
  1. The one-variable model adjusted R-squared value is 0.4836 and the two-variable model is 0.4768, making the one-variable model is lower than the two-variable method by about 1% when in percentage format. The one-variable model is better because the adjusted R-squared is higher. The higher the R-squared, the more variance is explained by the model.

  2. Summary paragraph Overall, we were able to able to look at all of the different measures of spread like the average, standard deviation, as well as the interquartile range for our two variables: bedrooms and bathrooms. The interquartile range was the best way to view the range of data since it was skewed. A boxplot and histogram were used to evaluate a different way to see the distribution of data. There were multiple outliers that could be identified through those graphs. Then both of the variables were graphed with prices to see if there was a correlation between the two variables. This allowed us to determine that there was a positive and weak relationship between bedrooms and price as well as bathrooms and price. In order to further evaluate the relationship between the two, a regression was performed and the R squared was compared between the two. A Q-Q residual plot was also plotted to see the normality of the data and consider if there needed to be any data transformation which we did not have to do since it was mostly linear. Another regression was then conducted to see the t-value and see which variable was truly significant or not. After looking at the t-value, p-value, and the R squared value it was decided that bedrooms were not significant while bathrooms were when it came to the price of Airbnb’s. At the very end, we were also able to determine that the one-variable model was better than the two-variable model since it described the data’s correlation more accurately.

By working in a group of three, we were all able to demonstrate our strengths in this project and help one another out with R studio if someone else had a question or wasn’t sure how to go about something. We were able to bounce ideas off each other about how to state certain values and the different interpretations of the data that we had. By the end of this project, we were all more comfortable using R studio as well as being able to interpret data that was presented in different formats.