Activity

In-Class R Coding Activity (Terry and Adam)

Step 1: Load the packages and data

  1. Load the tidyverse and dslabs packages.

  2. Load the murders dataset from the dslabs package.

  3. Import the state-level CSV dataset directly from the following URL using read_csv():

    https://raw.githubusercontent.com/plotly/datasets/master/2014_usa_states.csv

  4. Display the first six rows of each dataset

  5. How many rows and columns does each dataset contain? I can see 5 rows in he data set.

Step 2: Explore the datasets

  1. Use glimpse() to examine the structure of both datasets.

  2. Display the column names of each dataset using names().

  3. Identify the variable that represents the state in each dataset. The abbreivation variables represent the states. AL, AK, AZ, DC, CL, MD,MA

  4. Are the state variable names identical in both datasets? The two data sets share the same names

  5. Identify the variables that are common to both datasets. the names of states, population and total numbers, and regions

  6. Why is it important to inspect the datasets before joining them? It’s important so we can have the data we are looking for to prevent errors or mistakes.

3. Prepare the state names

The murders dataset has 51 observations, including Washington, D.C. The second dataset contains 52 rows, so not every row necessarily represents a state.

murders <- murders |>
  mutate(state = tolower(state))

scores <- ________________________

Step 4: Join the datasets

  1. Use left_join() to join the murders dataset with the state-level dataset.

  2. Use the state variable as the joining key.

  3. Store the joined dataset in a new object called combined_data.

  4. Display the first six rows of combined_data.

  5. How many rows does the joined dataset contain? my merged data contians 102 colums

    Export the final dataset

  6. Use write_csv() to export combined_data as a CSV file named combined_murders.csv

Step 5: Investigate unmatched states

  1. The murders dataset contains 51 observations, while the second dataset contains 52 rows. Why might the number of rows differ? They may differ becuase one data might have a colum that the other may not have.

  2. Use anti_join() to identify the states in murders that do not have a matching state in the second dataset.