dplyr is a core package in R as a part of the tidyverse that provides a set of “Verbs” which are used to transform data. dplyr is widely useful since it allows for the wrangling of data to become closer to plain English, which makes it more readable and more easily interpreted by people who are not sure what to look for.
In this demonstration, I will be using the Palmer Penguins dataset to show how to utilize each of the dplyr verbs one at a time to answer the question How do the three Palmer penguin species differ in size and bill shape?
| Verb | What it does |
|---|---|
select() |
Keeps the columns you need |
filter() |
Keeps the rows you need |
mutate() |
Creates new columns |
group_by() + summarize() |
Calculates results for each group |
arrange() |
Sorts rows |
The Palmer Penguins dataset consists of measurements of 344 penguins of 3 different species (Adelie, Chinstrap, and Gentoo). They were observed on 3 islands in the Palmer Archipelago between 2007 and 2009. The data was collected by Dr. Kristen Gorman and the Palmer Station Long Term Ecological Research program. Within the dataset, there are various physical measurements that allow us to compare the penguins and identify the differences between the species.
## Rows: 344
## Columns: 8
## $ species <fct> Adelie, Adelie, Adelie, Adelie, Adelie, Adelie, Adel…
## $ island <fct> Torgersen, Torgersen, Torgersen, Torgersen, Torgerse…
## $ bill_length_mm <dbl> 39.1, 39.5, 40.3, NA, 36.7, 39.3, 38.9, 39.2, 34.1, …
## $ bill_depth_mm <dbl> 18.7, 17.4, 18.0, NA, 19.3, 20.6, 17.8, 19.6, 18.1, …
## $ flipper_length_mm <int> 181, 186, 195, NA, 193, 190, 181, 195, 193, 190, 186…
## $ body_mass_g <int> 3750, 3800, 3250, NA, 3450, 3650, 3625, 4675, 3475, …
## $ sex <fct> male, female, female, NA, female, male, female, male…
## $ year <int> 2007, 2007, 2007, 2007, 2007, 2007, 2007, 2007, 2007…
The dataset contains 8 columns which contain different
characteristics of the penguins.
species: The identified species of the penguinisland: Where the penguin was foundbill_length_mm: Length of the penguin’s bill in
millimetersbill_depth_mm: Depth of the penguin’s bill in
millimetersflipper_length_mm: Length of the penguin’s flipper in
millimetersbody_mass_g: Mass of the penguin in gramssex: Identifies if the penguin was male or femaleyear: The year the penguin was identifiedThe pipe operator %>% is a tool in R that allows us
to connect commands together without having to pass the dataset into
commands several times. It takes the result of the data on the left and
passes it into the function on the right. One way to read these pipes is
to think of them as saying “and then”.
select() ColumnsThe select() command allows you to keep only the columns
from the data that you chose and drop everything else. Here, we do not
need flipper length or island to answer the question of How do the
three Palmer penguin species differ in size and bill shape?, so we
are able to select only the other columns.
penguin.size <- penguins %>%
select(species, sex, bill_length_mm, bill_depth_mm, body_mass_g)
head(penguin.size)Before using the select() command, there were 8 columns,
but after using the command, we can see only the 5 columns species, sex,
bill length, bill depth, and body mass. This allows us to focus only on
the measurements which are needed for our analysis. While the number of
columns have changed, the number of rows remains the same.
filter() RowsThe filter() command allows you to set a condition and
keep only rows that meet that condition, dropping the rest. This allows
us to work with individual observations in order to either analyze
certain categories or remove observations which are missing important
information.
Before we continue, we need to see if there are any penguins which are missing measurements. The is.na() function identifies penguins which do not have a value in the specified column. Combining the results with sum() allows us to turn this into a count. In this dataset, there are 2 penguins which are missing values for body mass and 11 which are missing values for sex. After removing these from the data, we are left with 333 penguins.
## [1] 2
## [1] 11
## [1] 333
Another way that we can use the filter() function is to
filter the data to view only specific categories. If we want to see only
Gentoo penguins, we can use filter() to keep only
observations where the species is Gentoo and exclude the others.
mutate() New ColumnsThe mutate() command allows us to create new columns
using data from existing columns in the dataset. In this example, we
want to create 3 new variables which will allow us to more easily
perform comparisons going forwards.
body_mass_kg: converts grams to kilograms, which are
easier to read.bill_ratio: bill length divided by bill depth. Higher
values mean a longer, thinner bill.heavy: TRUE if a penguin weighs more than 4.5 kg.penguin.size <- penguin.size %>%
mutate(body_mass_kg = body_mass_g / 1000,
bill_ratio = bill_length_mm / bill_depth_mm,
heavy = body_mass_kg > 4.5)
head(penguin.size)group_by() and summarize()The group_by() and summarize() commands are
helpful for comparing different categories within the same dataset.
group_by() organizes observations into groups based on a
selected variable, while summarize() calculates statistics
for each group. For this example, we want to group the penguins by
species and calculate the number of penguins, their average body mass,
their average bill ratio, and the proportion weighing more than 4.5
kilograms. This makes it much easier to compare the three species
without examining each penguin individually.
species.summary <- penguin.size %>%
group_by(species) %>%
summarize(penguins = n(),
avg.mass.kg = round(mean(body_mass_kg), 2),
avg.bill.ratio = round(mean(bill_ratio), 2),
share.heavy = round(mean(heavy), 2))
species.summaryWe are also able to use group_by() with multiple
variables in order to see more detailed comparisons. For this example,
we can group the penguins by both species and sex. This separates the
penguins into males and females of each species and then groups them by
size within those groups. This allows for us to see how weight patterns
differ between male and female penguins of each species.
penguin.size %>%
group_by(species, sex) %>%
summarize(penguins = n(),
avg.mass.kg = round(mean(body_mass_kg), 2))arrange() RowsAfter sorting the species by average body mass, we can see that the Gentoo penguins are the heaviest species. Adelie and Chinstrap penguins have similar average body masses which are much lower. By using desc(), we can tell R to sort from highest to lowest instead of lowest to highest, which is the default.
When we sort the species by their average bill ratio, we see a difference between Adelie and Chinstrap penguins. Although these two species have very similar average body masses, Chinstrap penguins have a higher average bill ratio, suggesting they have longer, thinner bills compared to Adelie penguins.
At this point, we are able to combine all of the operations into one continuous block of code. The pipe operator in combination with the verbs allows for us to do complex operations in a manner that remains readable. This is helpful in observing the penguin data because each step benefits from building upon the last.
penguins %>%
select(species, sex, bill_length_mm, bill_depth_mm, body_mass_g) %>%
filter(!is.na(body_mass_g),
!is.na(sex)) %>%
mutate(body_mass_kg = body_mass_g / 1000,
bill_ratio = bill_length_mm / bill_depth_mm,
heavy = body_mass_kg > 4.5) %>%
group_by(species) %>%
summarize(penguins = n(),
avg.mass.kg = round(mean(body_mass_kg), 2),
avg.bill.ratio = round(mean(bill_ratio), 2),
share.heavy = round(mean(heavy), 2)) %>%
arrange(desc(avg.mass.kg))Throughout this demonstration, we used the five core dplyr verbs to explore and analyze the Palmer Penguins dataset. We used select() to focus on specific columns, filter() to narrow down our observations, mutate() to create new variables, group_by() and summarize() to calculate statistics for different groups, and arrange() to organize our results. By combining these commands, we were able to identify several interesting differences between the three penguin species. Gentoo penguins had the highest average body mass, while Adelie and Chinstrap penguins had similar average weights but noticeable differences in their bill ratios. We also found that male penguins tended to weigh more than females within each species. Overall, this analysis demonstrates how the five core dplyr verbs can work together to make data easier to organize, explore, and interpret, allowing us to turn a larger dataset into meaningful insights with relatively simple code.
Horst, A. M., Hill, A. P., & Gorman, K. B. (2020). palmerpenguins: Palmer Archipelago (Antarctica) penguin data (R package). https://allisonhorst.github.io/palmerpenguins/