Author

Anastasiia Gmyrina

What is Pagination

APIs allow us to retrieve data from external sources, but large datasets may contain millions of records.

Pagination divides the results into smaller batches instead of returning everything at once.

For example, if an API contains 100,000 records and returns 1,000 per request, we need 100 requests to collect the full dataset.

Three common methods are:

  • Limit and offset
  • Page-number pagination
  • Cursor-based pagination

Limit and Offset Pagination

With limit-and-offset pagination, we control:

  • limit: How many records to retrieve.
  • offset: How many records to skip.

For example, a limit of 5 and an offset of 0 returns the first five records. Changing the offset to 5 retrieves the next five.

This method is easy to understand, but large offsets can become inefficient. It’s also important to sort the data consistently so that records don’t unexpectedly change positions between requests.

Show code
library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.3     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
Show code
library(jsonlite)

Attaching package: 'jsonlite'

The following object is masked from 'package:purrr':

    flatten
Show code
base_url <- "https://data.cityofnewyork.us/resource/h9gi-nx95.json"

page1 <- fromJSON(
  paste0(base_url, "?$limit=5&$offset=0&$order=collision_id"))

page2 <- fromJSON(
  paste0(base_url, "?$limit=5&$offset=5&$order=collision_id"))

bind_rows(page1, page2) %>%
  select(any_of(c("collision_id", "crash_date", "borough"))) %>%
  head(10)

Page-Number Pagination

The Rick and Morty API contains information about characters from the animated series, including names, species, and status.

Instead of telling the API how many records to skip, we request a specific page:

  • page=1: First page
  • page=2: Second page

The API returns up to 20 characters per page and provides information about the next and previous pages. This approach is convenient for websites where users browse results page by page. However, records can shift between pages when the underlying data changes.

Show code
page1 <- fromJSON("https://rickandmortyapi.com/api/character?page=1")

page2 <- fromJSON("https://rickandmortyapi.com/api/character?page=2")

characters <- bind_rows(page1$results, page2$results)

characters %>%
  select(id, name, status, species) %>%
  head(10)

Cursor-Based Pagination

Cursor-based pagination works like a bookmark. Instead of specifying a page number or how many records to skip, the API returns a cursor that tells us where to continue.

In this example, I first request five publications from 2024 using cursor=*, which starts at the beginning. The API returns the records along with a next_cursor value. I then use that cursor in a second request to retrieve the next five publications. Finally, I combine both batches into one data frame.

This method is especially useful for large datasets because we don’t need to calculate offsets. The API provides the information needed to continue retrieving records.

Show code
# First batch of 5 publications
page1 <- fromJSON(
  "https://api.openalex.org/works?filter=publication_year:2024&per_page=5&cursor=*"
)

# Get the cursor for the next batch
cursor <- page1$meta$next_cursor

# Request the next 5 publications
page2 <- fromJSON(
  paste0(
    "https://api.openalex.org/works?filter=publication_year:2024&per_page=5&cursor=",
    URLencode(cursor, reserved = TRUE)))

# Display the results
bind_rows(
  page1$results,
  page2$results
) %>%
  select(display_name, publication_year, cited_by_count)

Comparison of Pagination Methods

All three methods retrieve data in batches, but they identify the next batch differently.

  • Limit/offset: Skip a specified number of records.
  • Page number: Request the next numbered page.
  • Cursor: Continue using a token returned by the API.

The chart below illustrates the parameters used for three consecutive requests.

Show code
library(ggplot2)
library(tibble)

comparison <- tibble(
  method = rep(c("Limit & Offset", "Page Number", "Cursor"), each = 3),
  request = rep(1:3, times = 3),
  parameter = c(
    "offset = 0", "offset = 5", "offset = 10",
    "page = 1", "page = 2", "page = 3",
    "cursor = *", "next cursor", "next cursor"
  )
)

ggplot(comparison, aes(x = factor(request), y = method, fill = method)) +
  geom_tile(color = "white", linewidth = 3) +
  geom_text(aes(label = parameter), color = "white", size = 3.5) +
  labs(
    title = "How Pagination Methods Retrieve the Next Batch",
    x = "Request Number",
    y = NULL
  ) +
  theme_minimal() +
  theme(
    legend.position = "none",
    panel.grid = element_blank()
  )