Introduction

dplyr is a core package in R as a part of the tidyverse that provides a set of “Verbs” which are used to transform data. dplyr is widely useful since it allows for the wrangling of data to become closer to plain English, which makes it more readable and more easily interpreted by people who are not sure what to look for.

In this demonstration, I will be using the Palmer Penguins dataset to show how to utilize each of the dplyr verbs one at a time to answer the question How do the three Palmer penguin species differ in size and bill shape?

The Five Verbs

Verb What it does
select() Keeps the columns you need
filter() Keeps the rows you need
mutate() Creates new columns
group_by() + summarize() Calculates results for each group
arrange() Sorts rows


Setup

Load the Packages

library(dplyr)
library(palmerpenguins)

Meet the Data

The Palmer Penguins dataset consists of measurements of 344 penguins of 3 different species (Adelie, Chinstrap, and Gentoo). They were observed on 3 islands in the Palmer Archipelago between 2007 and 2009. The data was collected by Dr. Kristen Gorman and the Palmer Station Long Term Ecological Research program. Within the dataset, there are various physical measurements that allow us to compare the penguins and identify the differences between the species.

glimpse(penguins)
## Rows: 344
## Columns: 8
## $ species           <fct> Adelie, Adelie, Adelie, Adelie, Adelie, Adelie, Adel…
## $ island            <fct> Torgersen, Torgersen, Torgersen, Torgersen, Torgerse…
## $ bill_length_mm    <dbl> 39.1, 39.5, 40.3, NA, 36.7, 39.3, 38.9, 39.2, 34.1, …
## $ bill_depth_mm     <dbl> 18.7, 17.4, 18.0, NA, 19.3, 20.6, 17.8, 19.6, 18.1, …
## $ flipper_length_mm <int> 181, 186, 195, NA, 193, 190, 181, 195, 193, 190, 186…
## $ body_mass_g       <int> 3750, 3800, 3250, NA, 3450, 3650, 3625, 4675, 3475, …
## $ sex               <fct> male, female, female, NA, female, male, female, male…
## $ year              <int> 2007, 2007, 2007, 2007, 2007, 2007, 2007, 2007, 2007…

The dataset contains 8 columns which contain different characteristics of the penguins.

  • species: The identified species of the penguin
  • island: Where the penguin was found
  • bill_length_mm: Length of the penguin’s bill in millimeters
  • bill_depth_mm: Depth of the penguin’s bill in millimeters
  • flipper_length_mm: Length of the penguin’s flipper in millimeters
  • body_mass_g: Mass of the penguin in grams
  • sex: Identifies if the penguin was male or female
  • year: The year the penguin was identified


The Pipe Operator

The pipe operator %>% is a tool in R that allows us to connect commands together without having to pass the dataset into commands several times. It takes the result of the data on the left and passes it into the function on the right. One way to read these pipes is to think of them as saying “and then”.

Step 1: select() Columns

The select() command allows you to keep only the columns from the data that you chose and drop everything else. Here, we do not need flipper length or island to answer the question of How do the three Palmer penguin species differ in size and bill shape?, so we are able to select only the other columns.

penguin.size <- penguins %>% 
  select(species, sex, bill_length_mm, bill_depth_mm, body_mass_g)

head(penguin.size)

Before using the select() command, there were 8 columns, but after using the command, we can see only the 5 columns species, sex, bill length, bill depth, and body mass. This allows us to focus only on the measurements which are needed for our analysis. While the number of columns have changed, the number of rows remains the same.


Step 2: filter() Rows

The filter() command allows you to set a condition and keep only rows that meet that condition, dropping the rest. This allows us to work with individual observations in order to either analyze certain categories or remove observations which are missing important information.

Remove Missing Values

Before we continue, we need to see if there are any penguins which are missing measurements. The is.na() function identifies penguins which do not have a value in the specified column. Combining the results with sum() allows us to turn this into a count. In this dataset, there are 2 penguins which are missing values for body mass and 11 which are missing values for sex. After removing these from the data, we are left with 333 penguins.

sum(is.na(penguin.size$body_mass_g))
## [1] 2
sum(is.na(penguin.size$sex))
## [1] 11
penguin.size <- penguin.size %>% 
  filter(!is.na(body_mass_g), 
         !is.na(sex))

nrow(penguin.size)
## [1] 333

Filter by a Category

Another way that we can use the filter() function is to filter the data to view only specific categories. If we want to see only Gentoo penguins, we can use filter() to keep only observations where the species is Gentoo and exclude the others.

penguin.size %>% 
  filter(species == "Gentoo") %>% 
  head()

Step 3: mutate() New Columns

The mutate() command allows us to create new columns using data from existing columns in the dataset. In this example, we want to create 3 new variables which will allow us to more easily perform comparisons going forwards.

  • body_mass_kg: converts grams to kilograms, which are easier to read.
  • bill_ratio: bill length divided by bill depth. Higher values mean a longer, thinner bill.
  • heavy: TRUE if a penguin weighs more than 4.5 kg.
penguin.size <- penguin.size %>% 
  mutate(body_mass_kg = body_mass_g / 1000, 
         bill_ratio = bill_length_mm / bill_depth_mm, 
         heavy = body_mass_kg > 4.5)

head(penguin.size)


Step 4: group_by() and summarize()

The group_by() and summarize() commands are helpful for comparing different categories within the same dataset. group_by() organizes observations into groups based on a selected variable, while summarize() calculates statistics for each group. For this example, we want to group the penguins by species and calculate the number of penguins, their average body mass, their average bill ratio, and the proportion weighing more than 4.5 kilograms. This makes it much easier to compare the three species without examining each penguin individually.

Summarize by Species

species.summary <- penguin.size %>% 
  group_by(species) %>% 
  summarize(penguins = n(), 
            avg.mass.kg = round(mean(body_mass_kg), 2), 
            avg.bill.ratio = round(mean(bill_ratio), 2), 
            share.heavy = round(mean(heavy), 2))

species.summary

Group by More Than One Variable

We are also able to use group_by() with multiple variables in order to see more detailed comparisons. For this example, we can group the penguins by both species and sex. This separates the penguins into males and females of each species and then groups them by size within those groups. This allows for us to see how weight patterns differ between male and female penguins of each species.

penguin.size %>% 
  group_by(species, sex) %>% 
  summarize(penguins = n(), 
            avg.mass.kg = round(mean(body_mass_kg), 2))

Step 5: arrange() Rows

Heaviest Species

species.summary %>% 
  arrange(desc(avg.mass.kg))

After sorting the species by average body mass, we can see that the Gentoo penguins are the heaviest species. Adelie and Chinstrap penguins have similar average body masses which are much lower. By using desc(), we can tell R to sort from highest to lowest instead of lowest to highest, which is the default.

Longest, Thinnest Bills

species.summary %>% 
  arrange(desc(avg.bill.ratio))

When we sort the species by their average bill ratio, we see a difference between Adelie and Chinstrap penguins. Although these two species have very similar average body masses, Chinstrap penguins have a higher average bill ratio, suggesting they have longer, thinner bills compared to Adelie penguins.


Putting It All Together

At this point, we are able to combine all of the operations into one continuous block of code. The pipe operator in combination with the verbs allows for us to do complex operations in a manner that remains readable. This is helpful in observing the penguin data because each step benefits from building upon the last.

penguins %>% 
  select(species, sex, bill_length_mm, bill_depth_mm, body_mass_g) %>% 
  filter(!is.na(body_mass_g), 
         !is.na(sex)) %>% 
  mutate(body_mass_kg = body_mass_g / 1000, 
         bill_ratio = bill_length_mm / bill_depth_mm, 
         heavy = body_mass_kg > 4.5) %>% 
  group_by(species) %>% 
  summarize(penguins = n(), 
            avg.mass.kg = round(mean(body_mass_kg), 2), 
            avg.bill.ratio = round(mean(bill_ratio), 2), 
            share.heavy = round(mean(heavy), 2)) %>% 
  arrange(desc(avg.mass.kg))


Conclusion

Throughout this demonstration, we used the five core dplyr verbs to explore and analyze the Palmer Penguins dataset. We used select() to focus on specific columns, filter() to narrow down our observations, mutate() to create new variables, group_by() and summarize() to calculate statistics for different groups, and arrange() to organize our results. By combining these commands, we were able to identify several interesting differences between the three penguin species. Gentoo penguins had the highest average body mass, while Adelie and Chinstrap penguins had similar average weights but noticeable differences in their bill ratios. We also found that male penguins tended to weigh more than females within each species. Overall, this analysis demonstrates how the five core dplyr verbs can work together to make data easier to organize, explore, and interpret, allowing us to turn a larger dataset into meaningful insights with relatively simple code.


Works Cited

Horst, A. M., Hill, A. P., & Gorman, K. B. (2020). palmerpenguins: Palmer Archipelago (Antarctica) penguin data (R package). https://allisonhorst.github.io/palmerpenguins/