Introduction

The purpose of data visualization, via graphs, charts, diagrams, or tables is to uncover patterns in a dataset, and to ascertain if the patterns say something significant about the data. The visualization may be created as an exploratory tool for a particular dataset, to answer a research question, or provide evidence which supports or disproves an analytical theory and/or hypothesis. This code-through explores the importance of selecting the right visualization to best represent the dataset and how to spot a poor visual representation of data.


What Makes Data Visualization Good or Bad?

These days, with ready access to open data, and spreadsheet and presentation software that facilitates swift creation and access to data visualizations. And, although I like to say that numbers tell no lies, unfortunately, data visualizations may tell a few things, they may tell the truth, they may exaggerate the truth, or they may obstruct the truth (Lauer & O’Brien, 2020). It is also important to note that looks aren’t everything, a beautiful visualization is just as likely to be deceptive as a poorly constructed one.

The bottom line is that all data visualization should be reviewed with a trained and critical eye, i.e., do not assume that because you know the fundamentals about data visualization you are immune from deception.

The purpose of this code-through is to identify common ways misleading and/or deceptive data visuals are produced and provide some sound strategies for matching visualization to dataset for maximum clarity and impact. To move from the theoretical to practical application, a case study will be introduced to provide a roadmap for transforming a visualization which misrepresents the data to one which uncovers and enhances the patterns hidden in the data.


Nobody is Immune from Deception

Research has shown that even those well-studied in data visualization are as at much risk of being misled by bad visualizations as the general public (Lauer & O’Brien, 2020). Therefore, it behooves everyone to understand and adhere to best practices, while familiarizing themselves with common characteristics of deceptive data visualizations.


Common Deceptive Tactics

We don’t always know the reasons why–it could be “lack of training in statistics, …valu[ing] aesthetics over accuracy, …desir[ing] to show clear and simple trends [above all else]”, a miscalculation, faulty plotting or simply a direct attempt to “mislead their audience” (Lauer & O’Brien, 2020)–but the results are the same – faulty data visualization.

  • Distorting the Y-Axis: “One of the most common ways graphs misrepresent data, by distorting the scale. Zooming in on a small portion of the y-axis exaggerates a barely detectable difference between the things being compared. And it especially misleading with bar graphs since we assume the difference in the size of the bars is proportional to the values” (TED-Ed, 2017).
  • Manipulating the X-Axis: Also known as “cherry picking.” “A time range can be carefully chosen to exclude the impact of a major event right outside it. Picking specific data points can hide important changes in between” (TED-Ed, 2017).
  • Context and Omitted Data: “Even when there is nothing wrong with the graph itself, leaving out relevant data can give a misleading impression” (TED-Ed, 2017).
  • Message Reversal: This is also known as an inverted axis. (Lauer & O’Brien, 2020).
  • 3-D Beveling: A three dimensional representation of a data visualization, i.e., a bar or pie chart, which distorts the size of data (Lauer & O’Brien, 2020).
  • Arbitrary Sizing of Graphic Shapes: No direct correlations between the size of a visual element and the quantity it is representing leads to confusion and distorted interpretations of the data (Lauer & O’Brien, 2020).



CASE STUDY: Bad Visualization

For this case study we are analyzing an infographic produced on the Visual Capitalist Website, entitled, “Global Beer Consumption.” This visual is a Voronoi treemap that claims to represent “the top countries by total beer consumption” (Venditti, 2023). A treemap is known to represent parts of a whole, much like a pie chart. So the first question to ask is – does this dataset–total beer consumption by 25 countries plus an aggregated observation–represent parts of a whole, or something else?

Critical Assessment of Graphic:

To avoid violating copyright laws, I have provided a link to the original visualization Ranked: Which Countries Drink the Most Beer? and have attempted to replicate the original visualization minus the photographic element, fonts, and other aesthetic embellishments. Below is the step-by-step process I took to replicate the original.

# WRANGLE DATASET

voronoi_data <- beer %>%                # Create an object for dataset 
    select(country, ttl_cons_2021) %>%  # reorder the variables, first by country, then by total 2021 beer consumption 
    filter(!is.na(ttl_cons_2021)) %>%
  mutate(label = str_c(country, 
                       scales::comma(ttl_cons_2021), 
                       sep = "\n"))     # include labels that include country name and total 2021 beer consumption 
  
head(voronoi_data)                      # view the updated dataset


# SELECT OBSERVATIONS FOR DATASET

originalvisual_countries <- c(
  "Australia",
  "Brazil",
  "Canada",
  "Czech Republic",
  "Germany",
  "Japan",
  "Mexico",
  "Poland",
  "Romania",
  "Russia",
  "South Korea",
  "Spain",
  "United Kingdom",
  "United States of America",
  "Italy",
  "Thailand",
  "Ethiopia",
  "France",
  "South Africa",
  "India",
  "China",
  "Colombia",
  "Vietnam",
  "Ukraine",
  "Argentina"
)

originalvisual_countries
##  [1] "Australia"                "Brazil"                  
##  [3] "Canada"                   "Czech Republic"          
##  [5] "Germany"                  "Japan"                   
##  [7] "Mexico"                   "Poland"                  
##  [9] "Romania"                  "Russia"                  
## [11] "South Korea"              "Spain"                   
## [13] "United Kingdom"           "United States of America"
## [15] "Italy"                    "Thailand"                
## [17] "Ethiopia"                 "France"                  
## [19] "South Africa"             "India"                   
## [21] "China"                    "Colombia"                
## [23] "Vietnam"                  "Ukraine"                 
## [25] "Argentina"
# CREATE TREEMAP VISUALIZATION USING ALL THE DATA

beer_voronoi <- voronoiTreemap(
    data = voronoi_data,
    levels = c("label"),
    cell_size = "ttl_cons_2021",
    shape = "circle"
)

beer_colors <- c("#C97818")

drawTreemap(
    beer_voronoi,
    label_size = 1.5,
    label_color = "white",
    color_type = "categorical",
  color_palette = beer_colors, 
)

I realized after creating this visualization that the original visual did not include all countries in the dataset in their visual, only the top 25 and an aggregated observation labeled “Rest of the World.” Although I tried to locate the specific list of countries that were contained therein, I was unable to replicate it in its entirety. The original dataset accessed from Kirin Holdings only provided data for 53 countries, whereas the Visual Capitalist visualization included data from an additional 170 countries (and please note this is still not an accurate tally of world population which includes a total of 195 countries).

# CREATE AGGREGATED OBSERVATION 

original_voronoi_data <- beer %>%
  filter(!is.na(ttl_cons_2021)) %>%
  mutate(
    display_country = if_else(
      country %in% originalvisual_countries,
      country,
      "Rest of the World"
    )
  ) %>%
  group_by(display_country) %>%
  summarise(
    ttl_cons_2021 = sum(ttl_cons_2021),
    .groups = "drop"
  )

original_voronoi_data <- original_voronoi_data %>%
  mutate(
    label = str_c(
      display_country,
      scales::comma(ttl_cons_2021),
      sep = "\n"
    )
  )

original_voronoi_data

Next I constructed the weighted Voronoi treemap, similar to the original with one caveat, I created an aggregated object in RStudio for the remaining 28 countries in the dataset, not the 145 represented in the original.

# CREATE TREEMAP VISUALIZATION USING AGGREGATED OBSERVATION

original_beer_voronoi <- voronoiTreemap(
  data = original_voronoi_data,
  levels = c("label"),
  cell_size = "ttl_cons_2021",
  shape = "circle"
)

beer_colors <- c("#C97818")

drawTreemap(
  original_beer_voronoi,
  label_size = 1.5,
  color_type = "categorical",
  color_palette = beer_colors,
  label_color = "white"
)

Now that I have created, to the best of my ability, the original data visualization I will critically assess it. This includes an analysis of the labels, the numbers, the scale and the context.

Numbers & Labels: Faulty plotting

The Voronoi treemap was created to compare parts to a whole. So the first question to ask is this dataset a complete dataset, i.e., does it represent the entire world population? The answer is no. It is merely a subsection of “170 major countries and regions” (Global Beer Consumption by Country in 2020 | 2022 | KIRIN - Kirin Holdings Company, Limited, n.d.). Further, the labels for the smaller observations (i.e., countries) are difficult to read.

Context: Identifying Audience, Missing Data

Total beer consumption by country is typically of interest to beer manufacturers and distributors, not the general public.

# SORT TOTAL BEER CONSUMPTION BY COUNTRY

beer %>% 
  select(country, ttl_cons_2021) %>%      # reorder the variables, first by country, then by total 2021 beer consumption 
  filter(!is.na(ttl_cons_2021)) %>%       # keep observations where 2021 consumption is not missing
  arrange(desc(ttl_cons_2021))            # sort the data by 2021 consumption in descending order

Comparing this dataset to the dataset sorted by population…

# SORT BY POPULATION SIZE

beer %>% 
  select(country, pop_2021) %>%
  filter(!is.na(pop_2021)) %>%
  arrange(desc(pop_2021))

*Population numbers derived from https://data.worldbank.org/indicator.

What you see is that the return dataset sorted /by beer consumption/by country (descending order) is very similar to the dataset sorted by population size. Therefore it seems to communicate more about population than beer consumption. An apples-to-apples comparison of annual beer consumption is the solution. In other words a per-capita analysis is required to accurately compare countries by beer consumption.

# SORT BY PER_CAPITA

beer %>% 
  select(country, cons_vol_2021) %>%
  filter(!is.na(cons_vol_2021)) %>%
  arrange(desc(cons_vol_2021))

Sorting by per-capita beer consumption shows an interesting contrast to the population and total consumption data tables. It shows none of the countries represented in the top ten list of the other two tables, and further, several of the countries represented in the per-capita top ten happen to be from Eastern European countries.

Next we address the claim in the original article that beer consumption increased in 2021 compared to 2020 with no visual back-up. Only data from 2021 is represented. Luckily, I was able to locate the 2020 data from the same data source the creators of the original visualization accessed the 2021 data, on the Kirin Holdings website. Adding the 2020 numbers to the original dataset allows me to make that visual comparison and confirm or disprove the claim in the original article, of increased beer consumption from 2020 to 2021.

# ADD 2020 DATA TO ORIGINAL DATASET

beer %>%
  select(country, cons_vol_2020, cons_vol_2021) %>%
  filter(
    !is.na(cons_vol_2020) &
    !is.na(cons_vol_2021)
  )

Scale - Arbitrary Sizing of Graph Shapes

The Voronoi treemap represents beer consumption numbers by country with irregular shapes that compare areas proportionally to one another. Research has shown that it is harder for the human mind to accurately perceive differences in areas, especially those that are not only non-linear, they are irregular in general and specifically to one another (Perceptual Scaling of Map Symbols, 2007).



Tools to Ensure Quality Control

It can be challenging to narrow down the best visualization for any particular dataset. There are multiple issues that must be considered. Luckly big names in data visualization have forged a path for us to follow by providing best practices, guidelines for matching data to visualization, as well as a list of common ways that visualization goes wrong. Let’s take a closer look.

Steps to Good Visualization:

  • Ask yourself, is it good or bad Data? Does it fit the analysis you are trying to perform? Is it accurate, complete, valid, timely, unique, and consistent? (Jeremiah, 2025)
  • What type of data are you using, quantitative, qualitative or both? Are your values numeric, categorical, binary, or a combination of these? Knowing this will help you narrow down which visuals can be effectively matched to your dataset (Yau, 2013)
  • Is the visualization exploratory or explanatory? It is the difference between the visualization helping you understand the data better, or helping you communicate the findings of your research (LibGuides: Data Visualization: Best Practices, n.d.).
  • Who is your audience? Know your audience (Rougier et al., 2014) is it a group of peers, board review members, subject matter experts, or the general public. This should inform how much explanation should be provided i.e., the less knowledgeable the audience is of the subject matter the more explanation is needed, etc.
  • What aspects of the data are you trying to communicate? In other words, “Identify your message.” (LibGuides: Data Visualization: Best Practices, n.d.)
  • What coordinate system will you be using? Cartesian, Polar or Geographic? This will further narrow down which visualization are leveraged (Yau, 2013).
  • What is the context of your graphic presentation? Will it be presented in hard copy, or digitally via a monitor or a projection screen? This will inform your approach to the layout of the graphic. The best rule of thumb being, the farther away a visual the less information can be absorbed by the audience – i.e., as distance increases so should simplicity (Rougier et al., 2014).
  • Are you proving all the required information? Don’t take for granted that the audience can understand the visual without some type of textual support – “Captions are not optional” (Yau, 2013).
  • Are you relying too heavily on default settings? If you choose to keep a default setting understand its affect on the visual and rule out other options. In other words, do not blindly trust default settings (Rougier et al., 2014).
  • Are you using visual attributes to enhance the message? Use color, size, shape, and/or location thoughtfully and strategically (Rougier et al., 2014).
  • Is your graphic too cluttered? Avoid what is commonly referred to as “Chartjunk”, i.e., additional visualization attributes that do nothing to enhance the visual but their presence makes it more difficult to attend to the data being presented (Rougier et al., 2014)
  • “Are you compromising accuracy for aesthetic?” Message trumps beauty, [but] does it make sense?” You may need input from a objective third party to assist you with this process (Yau, 2013).
  • Are you using the right data visualization tool? There are many data visualization tools out there, are you leveraging the right data processing support for your analysis, i.e., R, Microsoft Office Suite, Matplotlib, Python, Inkscape, TikZ & PGF, GIMP, ImageMajic, etc. (Rougier et al., 2014).


Matching Data Types To Compatible Visualizations

Identify the most effective visualizations for your data by following general guidelines established by Nathan Yau, a well-known data statistician and visualization expert (2013):

Type of Data Recommended Visualization
Categorical Data bar graph, symbol plot
Parts of a Whole pie chart, stacked bar chart, Voronoi map
Subcategories treemap, mosaic plot
Time Series Data bar graph, line chart, dot plot, dot-bar graph, dumbbell chart
Cycles radial plot, calendar, rose diagram
Locations location map, connections
Regions choropleth map, contour map
Cartograms circular cartogram, diffusion-based cartogram
Multiple Variables scatterplots/loess curve, heat map, parallel coordinates plot
Distribution: Summary box plot, violin plot
Distribution: One Variable histogram, density plot
Distribution: Multiple Variables heat map, surface plot


SOLUTION: What is the data really trying to say?

Let’s return to the beer consumption visualization. How can we correct the desceptive elements of this visualization? Remember the goal is to transform the data into an effective visualization, one that “can be gauged by its simplicity, relevancy, and its ability to hold the user’s hand during their data discovery journey” (Coresignal, 2021).

Just looking over the dataset it is clear that Czech Republic’s total per-capita consumption is going to be an issue. It is 184.1 liters per-capita in 2021, almost double the amount by the next highest consumption rate by Austria which was 96.8. Therefore, including it in the data will most likely skew the visualization to the right, making all of the other countries too small to read visually. So to correct this, I first created a dumbbell chart of the remaining 24 countries, to provide a better comparative view between years…

# WRANGLE PER-CAPITA DATA

dumbbell_data <- beer %>%
  filter(
    country %in% originalvisual_countries,
    !is.na(cons_vol_2020),
    !is.na(cons_vol_2021)
  )%>%
  mutate(
    country = reorder(country, cons_vol_2021)
  )

main_data <- dumbbell_data %>%
  filter(country != "Czech Republic")

czech_data <- dumbbell_data %>%
  filter(country == "Czech Republic")


Next, I applied the new per-capita dataset to a compatible data visualization based on the guidelines listed above. We are looking at a time-series dataset, so after playing around with a bar chart for started, I landed on the dumbbell chart, as it highlights the differences between years instead of focusing on total per-capita numbers.

I then created a separate dumbbell chart for Czech Republic. To effectively communicate that it is a significantly larger observation than the remaining 24 countries, I provided a title and subtitle which identifies that this visualization is a truncated version of the original, with a separate scale, which is clearly marked.

Then we combined the two charts into one visualization.

# COMBINE 2 VISUALIZATIONS INTO 1

combined_plot <- main_plot / czech_plot +
  plot_layout(heights = c(8, 1.2))

combined_plot


What you get with the final visualization is a clear picture of comparative beer consumption, by country, by year (2020 to 2021). This starts to really tell you something about what the data is saying. First Czech Republic represents a significantly higher consumption rate than any of the other countries displayed in the dataset.It also shows where there is significant differences in beer consumption by year between countries, in fact, they are all over the map. I can’t see any particular pattern here.

Bottom line, we now have a chart that clearly communicates the data in a way that anyone, trained in data analysis or not, can understand. And it took significant more R code to get us there. And that may be the key. Pretty visuals are not that difficult to create but that doesn’t make them good data visualizations. Sometimes, as is the case here, the simplest visuals can take the most analysis and significantly more data wrangling and code construction to really make them come alive.



Works Cited

This code through references and cites the following sources: