Introduction

Network visualization is a method of data visualization that allows for researchers to conduct meaningful analysis of relational connections between observations in a data set. This code through introduces the visNetwork package in R to demonstrate how relational data can be transformed, organized, visualized, and examined in an interactive network. Using a data set of familiar consumer brands and their parent companies this tutorial demonstrates how network visualization can reveal connections between variables that maybe difficult to visualize in a traditional tables.


Content Overview

This code through demonstrates data manipulation and network visualization by explaining how to transform data into an interactive network, identify nodes and edges, organize relationships between variables, and create custom visual elements with interactive components allowing the user to explore relationships represented in the data.

The tutorial will demonstrate:
- Data Preparation: Using filter(), mutate(), case_when(), and transmute() to select, classify and organize variables.
- Combining Data Sets: Using left_join() and distinct() to incorporate new observations into the network.
- Network Creation: Visualizing relationships by creating nodes and edges using the visNetwork package.
- Customization Features: Using visGroups() to differentiate by the type of observation.
- Interactive Features: Using the visOptions() and visInteraction() functions to introduce selection menus, hovering, and navigation features.


Why Use Network Visualization?

Data scientists frequently work with data sets containing relationships. These connections are difficult to explore when viewing large data sets, even when organized into a clean table. Network visualization offers an accessible way to display and examine these relationships, identify patterns, and communicate meaningful connections. In this tutorials network visualization is applied to business relationships, but alternatively it can be used to examine relationships in any social networks, health care referral systems, organizational partnerships, financial transactions, or other interconnected systems. Learning to create interactive networks allows data scientists to explore relational data and communicate complex information more effectively.


Learning Objectives

This tutorial aims to teach how to:
- Prepare and organize relational data using functions from dplyr.
- Identify nodes and edges for network visualization.
- Use the visNetwork package to create and customize interactive networks.
- Incorporate interactive features to explore and interpret relationships within the network.



From Data Visualization to Interactive Network Visualization

Here, we’ll show how to use the dplyr and visNetwork packages to organize, manipulate, and visualize data. Functions from the dplyr package can be employed to filter observations, create new variables, summarize information, and combine data sets to prepare the data for analysis and visualization.

By combining traditional data manipulation with the visNetwork package this code through serves as an introduction to network visualization. The basic examples in the tutorial begin by demonstrating the basic structure of a network using a complete toy data set we will create to represent relationships between consumer brands and their parent companies. Readers will learn how to identify and create data frames for nodes and edges. The advanced examples demonstrate how to create a visualization with only a select subset of data, create and join new columns with additional categories into the existing data set, and how to customize and incorporate interactive user features for meaningful analysis and interpretation.


Building on Data Manipulation Techniques

This tutorial expands on a basic knowledge of data manipulation using the dplyr package, and data visualization taught in PAF513. The example builds on these skills to demonstrate the interactive data visualization capabilities of the visNetwork package.


Data Set Creation

Create a toy data set if the data is not already organized. This example data set was created from observational data on consumer brands and their parent company associations. To create the data set, use the assign <- function to assign and name to the new data frame. Create new variables, and use the rep() function to repeat a variable based on the number of relationships observations in the data set. Use the c() function to combines multiple observations into vectors. The condition stringsAsFactors = FALSE is used to preserve human readable names, otherwise character value variable names will be automatically translated into factors.

# Organize the variables into a data frame 
brands <- data.frame(
  
  parent = rep(c(     # The repeat function allows for multiple relationships
    "Alphabet Inc",
    "Amazon",
    "Apple Inc",
    "Berkshire Hathaway Inc",
    "Comcast Corporation",
    "General Motors", 
    "Johnson & Johnson",
    "Meta Platforms",
    "Microsoft Corporation",
    "Nestle",
    "PepsiCo",
    "Sony Group Corporation",
    "The Walt Disney Company",
    "Volkswagen Group",
    "Walmart Inc"
  ), times = c(
    3, 3, 2, 3, 2, 3, 2, 3, 3, 3, 3, 3, 3, 3, 2)),
  
  brand = c(        # Assign Brand Variables to their Associated Parent Company Variables 
    #Alphabet Inc
    "Google", "YouTube", "Waymo",
    #Amazon
    "Amazon Web Services", "Whole Foods Market", "Twitch",
    #Apple Inc
    "Beats Electronics", "Claris",
    #Berkshire Hathaway Inc
    "GEICO", "BNSF Railway", "Precision Castparts",
    #Comcast Corporation
    "NBCUniversal", "Sky Group",
    #General Motors
    "Chevrolet", "Cadillac", "GMC", 
    #Johnson & Johnson
    "DePuy Synthes", "Ethicon",
    #Meta Platforms
    "Facebook", "Instagram", "WhatsApp",
    #Microsoft Corporation
    "LinkedIn", "GitHub", "Xbox Game Studios",
    #Nestle
    "Nespresso", "Purina", "Gerber",
    #PepsiCo
    "Frito-Lay", "Quaker", "Gatorade",
    #Sony Group Corporation
    "Sony Pictures Entertainment", "Sony Music Entertainment", "Sony Interactive Entertainment",
    #The Walt Disney Company
    "ABC", "Pixar Animation Studios", "Marvel Entertainment",
    #Volkswagen Group
    "Audi", "Porsche", "Skoda Auto",
    #Walmart Inc
    "Sam's Club", "Great Value"
  ),
stringsAsFactors = FALSE)    # Prevents R from Converting Variables to Factors (Preserves Variables "Character Names") 

Data Note: This data set was manually compiled using publicly available corporate information for educational purposes. It represents relationships/associations between the selected brands and parent companies, and is not an exhaustive database.


Exploring the Data Set

In order to create meaningful visualizations and analysis, it is useful to view the data set to examine the structure prior to creating a network visualization. The head() function displays the first few observations contained in a data set. The dim() function identifies the number of rows and columns. The count() function from dplyr summarizes the number of associations between the parent companies.

# View the data set
knitr::kable(
  head(brands, 10),                       #Edit the number to choose the number of observations to show
  caption = "Brands and their Parent Companies"
)
Brands and their Parent Companies
parent brand
Alphabet Inc Google
Alphabet Inc YouTube
Alphabet Inc Waymo
Amazon Amazon Web Services
Amazon Whole Foods Market
Amazon Twitch
Apple Inc Beats Electronics
Apple Inc Claris
Berkshire Hathaway Inc GEICO
Berkshire Hathaway Inc BNSF Railway
# View data set dimensions
dim(brands)
## [1] 41  2
# Use count() to view the number of relationships observed for each individual variable in a selected column  
brands %>%   
  count(parent, name = "# of brands")


Basic Examples: Prepare the data frame

This example shows how to create a data frame containing the full data set, without additions or manipulation. Creating a network requires the programmer to identify nodes and edges.

Nodes identify individual variables, represented by parent companies and their associated brand relationships. In this example, the unique() function is used first to identify relationships between variables. In this example brands are assigned to their associated parent companies to prevent duplicates.

Edges create the connections between nodes, in this example each edge connects a parent company to the related brand variables. The transmute() function can be applied to assist in the creation of edges using the column names and the conditions (to) & (from) to represent relationships.

# Identify unique relationships between parent companies and brand variables 
parents <- unique(brands$parent)

# Create Nodes Data Frame
nodes <- data.frame(id = c(parents, brands$brand),
                    label = c(parents, brands$brand),
                    group = c(
                      rep("Parent Company", length(parents)), 
                      rep("Brand", nrow(brands))
                    )
)

# Use Transmute to Create Edges Data Frame

edges <- brands %>%
  transmute(
    from = parent,                 
    to = brand
  )


Basic Network Visualization Example:

**Use the visNetwork package to create a network visualization, that is inclusive of the entire data set. This example uses the nodes and edges identified above to include all observed parent companies and brand relationships contained in the data set. Conditions beyond this basic visualization allow further customization and user interactivity.

#Create an Exhaustive network Containing the Full Data Set

visNetwork(nodes, 
           edges, 
           height = "650px", 
           width = "100%", 
           main = "Corporate Brand Relationships"
)



Advanced Network Visualization Examples: Use Data Manipulation to Create an Exclusive Network Visualization

This example demonstrates how to choose a single selected variable and the related observations in cases where it is meaningful to examine a smaller selection of the data set. In this example only the parent company “Meta Platforms” and the associated brands “facebook”, “Instagram”, and “Whats App” are selected.

First create the nodes for the exclusive network visualization. Use the filter() function to filter for the targeted observation. Only variables with a relationship to the selected variable will populate in the visualization.

# Select the variable parent company 
meta <- brands %>%
  filter(parent == "Meta Platforms")  

# Preview 
meta

Then, define the nodes. Use the data.frame() function and assign operator <- to create and name the new data frame. The example uses the c() function to combine the parent company and associate brands. The conditions id and label allow the programmer to control the display name for each variable in the visualization. Use the assigned data frame name to preview the selection.

# Create the Nodes 
meta.nodes <- data.frame(
  id = c("Meta Platforms", meta$brand),   
  label = c("Meta Platforms", meta$brand)
)
# View the nodes for the selected data frame
meta.nodes

Next, create the edges using the transmute() function to identify which relationships will be represented. Define these connections between nodes with the conditions (from) and (to). Use the assigned data frame name to preview the selection.

# Create the Edges
meta.edges <- meta %>%
  transmute(
    from = parent,
    to = brand
  )
# View the edges for the selected data frame
meta.edges

Then use the visNetwork package to create a network visualization from the selected data frame. This example below allows users to explore relationships contained within the selected data frame, containing only observations associated with “Meta Platforms”. Only the parent company “Meta Platforms” and the associated brands are defined by meta.nodes and meta.edges above.

# Create a Graphic to Visualize the Relationships Contained in the Filtered Network 
visNetwork(
  meta.nodes,
  meta.edges,
  height = "400px",
  width = "100%"
)



##Advanced Examples: Adding A Column Use deplyr() to manipulate and edit the data set. The mutate() function allows programmers to add a new column allowing for the addition of new variable associations. The example code adds a categorical variable for corporate industry.

# Use mutate to Add New Variables   
brands <- brands %>%         # Use Assign to Name and Save the New Data Frame 
  mutate(                           #Use the mutate() Function to Add a New Column of Observations 
    industry = case_when(
      parent %in% c(
        "Alphabet Inc",
        "Apple Inc",
        "Meta Platforms",
        "Microsoft Corporations"
      ) ~ "Technology",
      
      parent == "Amazon" ~ "Technology & Retail",  # Use Exactly Equal to When there is Only One Observation
      
      parent %in% c(            # Use Combine When the Additional Variable is Assoicated with Multiple Observations 
        "The Walt Disney Company",
        "Comcast Corporation",
        "Sony Group Corporation"
        ) ~ "Media & Entertainment",
      
      parent %in% c(
        "Volkswagen Group",
        "General Motors"
      ) ~ "Automotive",
      
      parent %in% c(
        "Nestle", 
        "PepsiCo"
      ) ~ "Food & Beverage",
      
      parent == "Johnson & Johnson" ~ "Healthcare",  # Use Exactly Equal to When there is Only One Observation
      parent == "JPMorgan Chase Co" ~ "Finance",
      parent == "Birkshire Hathaway Inc" ~ "Conglomerate",
      parent == "Walmart Inc" ~ "Retail",
      
      # Set a Default Category for Uncategorized Brands 
      TRUE ~ "Other"
    )
  )
# Preview the New Data Frame with the Addition of the New Column
head(brands, 10)
# Count the Number of Observations Contained in Industry Classifications  
brands %>%
  count(industry)

Data Note: Industry categories were assigned for this tutorial.


Advanced Customization Example: Interactive Data Visualization

Most notably, the visNetwork package can be used to customize the visual for meaningful analysis. The left_join function is applied to add the newly created categorical variable of “Industry” to the nodes. Use the left_join() function to join the data set and the new data frame created above. In this example the visNetwork and dplyr packages are employed in tandem to customize and create more advanced interactive features. This example demonstrates how visGroup(), visOptions(), mutate() and case_when() functions each contribute to interactive visualization.

# Use the left_join() Function to Combine the New Data Frame with Original Data Set    
nodes <- nodes %>%   
  left_join(
    brands %>%
      distinct(parent, industry),
    by = c("id" = "parent")
  ) %>%
  left_join(
    brands %>%
      select(brand, parent, industry) %>%
      rename(brand.industry = industry),
    by = c("id" = "brand")
  ) %>%
  mutate(
    industry = coalesce(industry, brand.industry),
    # Customize the Size of the Nodes by Type                
    size = ifelse(            
      group == "Parent Company", 
      35, 
      15
    ),
# Add the Interactive Categorical Variables 
title = ifelse(
  group == "Parent Company",
  paste0(
    "<b>", label, "</b>", 
    "<br>Industry:", industry
  ),
  paste0(
    "<b>", label, "</b>", 
    "<br>Parent Company:", parent,
    "<br>Industry:", industry 
  )
)
) %>%
select (-parent, -brand.industry)

#Use visNetwork() and visGroups() to Create the Interactive Network 
visNetwork(nodes, 
           edges, 
           height = "650px", 
           width = "100%", 
           main = "Corportate Brand Relationships"
) %>%

  visGroups(                       # Select the Variable Group to Customize 
    groupname = "Parent Company",  # Name the Group/Selected Variables  
    color = "#177E89",             # Set a Unique Color for this Group
    shape = "dot"                  # Change the Shape for this Group
  ) %>%
  
  visGroups(
    groupname = "Brand",
    color = "#E9B44C",           
    shape = "dot"
  ) %>%
  
  # Use the visOptions() function to Add Interactive Selection 
  visOptions(
    highlightNearest = TRUE,
    nodesIdSelection = list(
      enabled = TRUE,
      main = "Select Brand Name:"      # Name the Drop Down Box
    ),
    selectedBy = list(
      variable = "industry",
      main = "Select an Industry:"    # Name the Drop Down Box
    )
  ) %>%
  
  # Use visInteraction() to Add Hovering and Navigation 
  visInteraction(
    hover = TRUE,
    navigationButtons = TRUE
  )




Further Resources

Learn more about the data set, how to create data frames, and the dplyr and visNetwork packahes and techniques with the following:


Works Cited

This code through references and cites the following sources: