Source Data

The following report analyzes tax parcel data from Syracuse, New York (USA).

View the “Data Dictionary” here: Syracuse City Tax Parcel Data



Importing the Data

The following code imports the Syracuse, NY tax parcel data using a URL.

url <- paste0("https://raw.githubusercontent.com/DS4PS/Data",
              "-Science-Class/master/DATA/syr_parcels.csv")

dat <- read.csv(url, 
                strings = FALSE)



Previewing the Data

There are several exploratory functions to better understand our new dataset.

We can inspect the first 5 rows of these data using function head().


head(dat, 5)              # Preview a dataset with 'head()'


Listing All Variables

Functions names() or colnames() will print all variable names in a dataset.


names(dat)                # List all variables with 'names()'
##  [1] "tax_id"       "neighborhood" "stnum"        "stname"       "zip"         
##  [6] "owner"        "frontfeet"    "depth"        "sqft"         "acres"       
## [11] "yearbuilt"    "age"          "age_range"    "land_use"     "units"       
## [16] "residential"  "rental"       "vacantbuil"   "assessedla"   "assessedva"  
## [21] "tax.exempt"   "countytxbl"   "schooltxbl"   "citytaxabl"   "star"        
## [26] "amtdelinqu"   "taxyrsdeli"   "totint"       "overduewater"


Previewing Specific Variables

We can also inspect the values of a variable by extracting it with $.

The extracted variable is called a “vector”.


head(dat$owner, 10)       # Preview a variable, or "vector"
##  [1] "CLARMIN BUILDERS ONON COR" "JOHNSTON LEE R"           
##  [3] "CHRISTO CRAIG S"           "HAWKINS FARMS INC"        
##  [5] "PETERS LYNNETTE"           "MITCHELL LOTAN G"         
##  [7] "WHALEN GIOVANNA A"         "BERGH GARY D"             
##  [9] "CITY OF SYRACUSE TD"       "DOUGHERTY ROBERT K JR"


Listing Unique Values

Function unique() helps us determine what values exist in a variable.


unique(dat$land_use)      # Print all possible values with 'unique()'
##  [1] "Vacant Land"        "Single Family"      "Commercial"        
##  [4] "Parking"            "Two Family"         "Three Family"      
##  [7] "Apartment"          "Schools"            "Parks"             
## [10] "Multiple Residence" "Cemetery"           "Religious"         
## [13] "Recreation"         "Community Services" "Utilities"         
## [16] "Industrial"


Examining Data Structure

Function str() provides an overview of total rows and columns (dimensions), variable classes, and a preview of values.


str(object = dat,
    vec.len = 2)          # Examine data structure with 'str()'
## 'data.frame':    41502 obs. of  29 variables:
##  $ tax_id      : int  1393130501 1393130500 1437100600 1425100900 1425101000 ...
##  $ neighborhood: chr  "South Valley" "South Valley" ...
##  $ stnum       : chr  "2655" "2635" ...
##  $ stname      : chr  "VALLEY DR" "VALLEY DR" ...
##  $ zip         : chr  "13215" "13120" ...
##  $ owner       : chr  "CLARMIN BUILDERS ONON COR" "JOHNSTON LEE R" ...
##  $ frontfeet   : num  67.2 104.8 ...
##  $ depth       : num  50 46.5 ...
##  $ sqft        : num  2149 6370 ...
##  $ acres       : num  0.0493 0.1462 ...
##  $ yearbuilt   : int  NA 1925 1957 1958 1965 ...
##  $ age         : int  NA 90 58 57 50 ...
##  $ age_range   : chr  NA "81-90" ...
##  $ land_use    : chr  "Vacant Land" "Single Family" ...
##  $ units       : int  0 0 0 0 0 ...
##  $ residential : logi  FALSE TRUE TRUE ...
##  $ rental      : logi  FALSE FALSE FALSE ...
##  $ vacantbuil  : logi  FALSE FALSE FALSE ...
##  $ assessedla  : int  475 10800 20200 18000 18000 ...
##  $ assessedva  : int  500 69300 88300 70500 74000 ...
##  $ tax.exempt  : logi  TRUE FALSE FALSE ...
##  $ countytxbl  : int  500 69300 88300 70500 74000 ...
##  $ schooltxbl  : int  500 69300 88300 70500 74000 ...
##  $ citytaxabl  : int  500 69300 88300 70500 74000 ...
##  $ star        : logi  NA TRUE TRUE ...
##  $ amtdelinqu  : num  0 0 0 0 0 ...
##  $ taxyrsdeli  : int  0 0 0 0 0 ...
##  $ totint      : num  0 0 0 0 0 ...
##  $ overduewater: num  0 178 ...



Questions & Solutions

Instructions: Provide the code for each solution in the following “chunks”.

Remember to modify the text to show your answer in human-readable terms.


Question 1: Total Parcels

Question: How many tax parcels are in Syracuse, NY?

Answer: There are [X] tax parcels in Syracuse, NY.


# Use an exploratory function like 'dim()', 'nrow()', or 'str()'


Question 2: Total Acres

Question: How many acres of land are in Syracuse, NY?

Answer: There are [X] acres of land in Syracuse, NY.


# Pass a numeric variable to function 'sum()', with argument 'na.rm = TRUE'


Question 3: Vacant Buildings

Question: How many vacant buildings are there in Syracuse, NY?

Answer: There are [X] vacant buildings in Syracuse, NY.


# Pass a numeric variable to function 'sum()', with argument 'na.rm = TRUE'

Question 4: Tax-Exempt Parcels

Question: What proportion of parcels are tax-exempt?

Answer: [X]% of parcels are tax-exempt.


# Pass a logical ('TRUE' or 'FALSE') variable to function 'mean()', with argument 'na.rm = TRUE'


Question 5: Neighborhoods & Parcels

Question: Which neighborhood contains the most tax parcels?

Answer: [X] contains the most tax parcels.


# Pass the appropriate variable to function 'table()'

# Optional: Use additional functions to narrow your results


Question 6: Neighborhoods & Vacant Lots

Question: Which neighborhood contains the most vacant lots?

Answer: [X] contains the most vacant lots.


# Pass two variables to function 'table()', separated by a comma

# (Optional) use additional functions to narrow your results



————

DELETE THIS LINE & ALL LINES BELOW BEFORE SUBMITTING

————



Tips & Tricks

The following tips and tricks are essential to understand in order to complete this assignment.


Extracting Variables

Reference variables in R by using the dataset name and variable name, separated by the $ operator.

summary(dat$acres)      # Extract variable 'acres' from dataset 'dat'


Adding Up Values

Use function sum() with a numeric vector to return the sum of all of its values.

sum(c(10, 20, 5))       # Add up values 10, 20, and 5
## [1] 35
sum(dat$sqft)           # Add up variable 'sqft' from dataset 'dat'
## [1] 544956951


Adding Up True Values

Function sum() with logical values, i.e. TRUE and FALSE values, will add up all instances of TRUE.

x <- c(TRUE, TRUE, 
       FALSE, FALSE, 
       FALSE, FALSE)    # Creating a vector of logical values: 'x'

sum(x)                  # Determining the total 'TRUE' values
## [1] 2


Proportions of True Values

Function mean() with logical values will provide the total proportion of TRUE values.

x <- c(TRUE, TRUE, 
       FALSE, FALSE, 
       FALSE, FALSE)    # Creating a vector of logical values: 'x'

mean(x)                 # Determining the proportion of 'TRUE' values
## [1] 0.3333333


Sums & Proportions with Missing Values

R wants you to know if sum(), mean() etc. include missing or NA values.

If a value is missing (NA), these functions will return NA.

y <- c(2, 3, NA, 5)     # Creating a vector with a missing ('NA') value: 'y'

sum(y)                  # Attempting to add up the values in 'y'
## [1] NA


Allowing Calculations with Missing Values

Use argument na.rm = TRUE to allow functions to ignore missing (NA) values.

y <- c(2, 3, NA, 5)     # Creating a vector with a missing ('NA') value: 'y'

sum(y, na.rm = TRUE)    # Add up the values in 'y' while ignoring missing values
## [1] 10



How to Submit

Use the following instructions to submit your assignment, which may vary depending on your course’s platform.


Knitting to HTML

When you have completed your assignment, click the “Knit” button to render your .RMD file into a .HTML report.


Special Instructions

Perform the following depending on your course’s platform:

  • Canvas: Upload both your .RMD and .HTML files to the appropriate link
  • Blackboard or iCollege: Compress your .RMD and .HTML files in a .ZIP file and upload to the appropriate link

.HTML files are preferred but not allowed by all platforms.


Before You Submit

Remember to ensure the following before submitting your assignment.

  1. Name your files using this format: Lab-##-LastName.rmd and Lab-##-LastName.html
  2. Show both the solution for your code and write out your answers in the body text
  3. Do not show excessive output; truncate your output, e.g. with function head()
  4. Follow appropriate styling conventions, e.g. spaces after commas, etc.
  5. Above all, ensure that your conventions are consistent

See Google’s R Style Guide for examples of common conventions.



Common Knitting Issues

.RMD files are knit into .HTML and other formats procedural, or line-by-line.

  • An error in code when knitting will halt the process; error messages will tell you the specific line with the error
  • Certain functions like install.packages() or setwd() are bound to cause errors in knitting
  • Altering a dataset or variable in one chunk will affect their use in all later chunks
  • If an object is “not found”, make sure it was created or loaded with library() in a previous chunk

If All Else Fails: If you cannot determine and fix the errors in a code chunk that’s preventing you from knitting your document, add eval = FALSE inside the brackets of {r} at the beginning of a chunk to ensure that R does not attempt to evaluate it, that is: {r eval = FALSE}. This will prevent an erroneous chunk of code from halting the knitting process.