Network visualization is a method of data visualization that allows
for researchers to conduct meaningful analysis of relational connections
between observations in a data set. This code through introduces the
visNetwork package in R to demonstrate how relational data
can be transformed, organized, visualized, and examined in an
interactive network. Using a data set of familiar consumer brands and
their parent companies this tutorial demonstrates how network
visualization can reveal connections between variables that maybe
difficult to visualize in a traditional tables.
This code through demonstrates data manipulation and network visualization by explaining how to transform data into an interactive network, identify nodes and edges, organize relationships between variables, and create custom visual elements with interactive components allowing the user to explore relationships represented in the data.
The tutorial will demonstrate:
- Data Preparation: Using
filter(), mutate(), case_when(),
and transmute() to select, classify and organize variables.
- Combining Data Sets: Using left_join() and
distinct() to incorporate new observations into the
network.
- Network Creation: Visualizing relationships by creating
nodes and edges using the visNetwork package.
-
Customization Features: Using visGroups() to differentiate
by the type of observation.
- Interactive Features: Using the
visOptions() and visInteraction() functions to
introduce selection menus, hovering, and navigation features.
Data scientists frequently work with data sets containing relationships. These connections are difficult to explore when viewing large data sets, even when organized into a clean table. Network visualization offers an accessible way to display and examine these relationships, identify patterns, and communicate meaningful connections. In this tutorials network visualization is applied to business relationships, but alternatively it can be used to examine relationships in any social networks, health care referral systems, organizational partnerships, financial transactions, or other interconnected systems. Learning to create interactive networks allows data scientists to explore relational data and communicate complex information more effectively.
This tutorial aims to teach how to:
- Prepare and organize
relational data using functions from dplyr.
- Identify
nodes and edges for network visualization.
- Use the
visNetwork package to create and customize interactive
networks.
- Incorporate interactive features to explore and
interpret relationships within the network.
Here, we’ll show how to use the dplyr and
visNetwork packages to organize, manipulate, and visualize
data. Functions from the dplyr package can be employed to
filter observations, create new variables, summarize information, and
combine data sets to prepare the data for analysis and
visualization.
By combining traditional data manipulation with the
visNetwork package this code through serves as an
introduction to network visualization. The basic examples in the
tutorial begin by demonstrating the basic structure of a network using a
complete toy data set we will create to represent relationships between
consumer brands and their parent companies. Readers will learn how to
identify and create data frames for nodes and
edges. The advanced examples demonstrate how to create a
visualization with only a select subset of data, create and join new
columns with additional categories into the existing data set, and how
to customize and incorporate interactive user features for meaningful
analysis and interpretation.
This tutorial expands on a basic knowledge of data manipulation using
the dplyr package, and data visualization taught in PAF513.
The example builds on these skills to demonstrate the interactive data
visualization capabilities of the visNetwork package.
Create a toy data set if the data is not already organized. This
example data set was created from observational data on consumer brands
and their parent company associations. To create the data set, use the
assign <- function to assign and name to the new data
frame. Create new variables, and use the rep() function to
repeat a variable based on the number of relationships observations in
the data set. Use the c() function to combines multiple
observations into vectors. The condition
stringsAsFactors = FALSE is used to preserve human readable
names, otherwise character value variable names will be automatically
translated into factors.
# Organize the variables into a data frame
brands <- data.frame(
parent = rep(c( # The repeat function allows for multiple relationships
"Alphabet Inc",
"Amazon",
"Apple Inc",
"Berkshire Hathaway Inc",
"Comcast Corporation",
"General Motors",
"Johnson & Johnson",
"Meta Platforms",
"Microsoft Corporation",
"Nestle",
"PepsiCo",
"Sony Group Corporation",
"The Walt Disney Company",
"Volkswagen Group",
"Walmart Inc"
), times = c(
3, 3, 2, 3, 2, 3, 2, 3, 3, 3, 3, 3, 3, 3, 2)),
brand = c( # Assign Brand Variables to their Associated Parent Company Variables
#Alphabet Inc
"Google", "YouTube", "Waymo",
#Amazon
"Amazon Web Services", "Whole Foods Market", "Twitch",
#Apple Inc
"Beats Electronics", "Claris",
#Berkshire Hathaway Inc
"GEICO", "BNSF Railway", "Precision Castparts",
#Comcast Corporation
"NBCUniversal", "Sky Group",
#General Motors
"Chevrolet", "Cadillac", "GMC",
#Johnson & Johnson
"DePuy Synthes", "Ethicon",
#Meta Platforms
"Facebook", "Instagram", "WhatsApp",
#Microsoft Corporation
"LinkedIn", "GitHub", "Xbox Game Studios",
#Nestle
"Nespresso", "Purina", "Gerber",
#PepsiCo
"Frito-Lay", "Quaker", "Gatorade",
#Sony Group Corporation
"Sony Pictures Entertainment", "Sony Music Entertainment", "Sony Interactive Entertainment",
#The Walt Disney Company
"ABC", "Pixar Animation Studios", "Marvel Entertainment",
#Volkswagen Group
"Audi", "Porsche", "Skoda Auto",
#Walmart Inc
"Sam's Club", "Great Value"
),
stringsAsFactors = FALSE) # Prevents R from Converting Variables to Factors (Preserves Variables "Character Names") Data Note: This data set was manually compiled using publicly available corporate information for educational purposes. It represents relationships/associations between the selected brands and parent companies, and is not an exhaustive database.
In order to create meaningful visualizations and analysis, it is
useful to view the data set to examine the structure prior to creating a
network visualization. The head() function displays the
first few observations contained in a data set. The dim()
function identifies the number of rows and columns. The
count() function from dplyr summarizes the
number of associations between the parent companies.
# View the data set
knitr::kable(
head(brands, 10), #Edit the number to choose the number of observations to show
caption = "Brands and their Parent Companies"
)| parent | brand |
|---|---|
| Alphabet Inc | |
| Alphabet Inc | YouTube |
| Alphabet Inc | Waymo |
| Amazon | Amazon Web Services |
| Amazon | Whole Foods Market |
| Amazon | Twitch |
| Apple Inc | Beats Electronics |
| Apple Inc | Claris |
| Berkshire Hathaway Inc | GEICO |
| Berkshire Hathaway Inc | BNSF Railway |
## [1] 41 2
# Use count() to view the number of relationships observed for each individual variable in a selected column
brands %>%
count(parent, name = "# of brands")This example shows how to create a data frame containing the full data set, without additions or manipulation. Creating a network requires the programmer to identify nodes and edges.
Nodes identify individual variables, represented by
parent companies and their associated brand relationships. In this
example, the unique() function is used first to identify
relationships between variables. In this example brands are assigned to
their associated parent companies to prevent duplicates.
Edges create the connections between nodes, in this
example each edge connects a parent company to the related brand
variables. The transmute() function can be applied to
assist in the creation of edges using the column names and the
conditions (to) & (from) to represent
relationships.
# Identify unique relationships between parent companies and brand variables
parents <- unique(brands$parent)
# Create Nodes Data Frame
nodes <- data.frame(id = c(parents, brands$brand),
label = c(parents, brands$brand),
group = c(
rep("Parent Company", length(parents)),
rep("Brand", nrow(brands))
)
)
# Use Transmute to Create Edges Data Frame
edges <- brands %>%
transmute(
from = parent,
to = brand
)**Use the visNetwork package to create a network
visualization, that is inclusive of the entire data set. This example
uses the nodes and edges identified above to include all observed parent
companies and brand relationships contained in the data set. Conditions
beyond this basic visualization allow further customization and user
interactivity.
#Create an Exhaustive network Containing the Full Data Set
visNetwork(nodes,
edges,
height = "650px",
width = "100%",
main = "Corporate Brand Relationships"
)This example demonstrates how to choose a single selected variable and the related observations in cases where it is meaningful to examine a smaller selection of the data set. In this example only the parent company “Meta Platforms” and the associated brands “facebook”, “Instagram”, and “Whats App” are selected.
First create the nodes for the exclusive network visualization. Use
the filter() function to filter for the targeted
observation. Only variables with a relationship to the selected variable
will populate in the visualization.
# Select the variable parent company
meta <- brands %>%
filter(parent == "Meta Platforms")
# Preview
metaThen, define the nodes. Use the data.frame() function
and assign operator <- to create and name the new data
frame. The example uses the c() function to combine the
parent company and associate brands. The conditions id and label allow
the programmer to control the display name for each variable in the
visualization. Use the assigned data frame name to preview the
selection.
# Create the Nodes
meta.nodes <- data.frame(
id = c("Meta Platforms", meta$brand),
label = c("Meta Platforms", meta$brand)
)
# View the nodes for the selected data frame
meta.nodesNext, create the edges using the transmute() function to
identify which relationships will be represented. Define these
connections between nodes with the conditions (from) and
(to). Use the assigned data frame name to
preview the selection.
# Create the Edges
meta.edges <- meta %>%
transmute(
from = parent,
to = brand
)
# View the edges for the selected data frame
meta.edgesThen use the visNetwork package to create a network
visualization from the selected data frame. This example below allows
users to explore relationships contained within the selected data frame,
containing only observations associated with “Meta Platforms”. Only the
parent company “Meta Platforms” and the associated brands are defined by
meta.nodes and meta.edges above.
# Create a Graphic to Visualize the Relationships Contained in the Filtered Network
visNetwork(
meta.nodes,
meta.edges,
height = "400px",
width = "100%"
)
##Advanced Examples: Adding A Column Use
deplyr() to manipulate and edit the data set. The
mutate() function allows programmers to add a new column
allowing for the addition of new variable associations. The example code
adds a categorical variable for corporate industry.
# Use mutate to Add New Variables
brands <- brands %>% # Use Assign to Name and Save the New Data Frame
mutate( #Use the mutate() Function to Add a New Column of Observations
industry = case_when(
parent %in% c(
"Alphabet Inc",
"Apple Inc",
"Meta Platforms",
"Microsoft Corporations"
) ~ "Technology",
parent == "Amazon" ~ "Technology & Retail", # Use Exactly Equal to When there is Only One Observation
parent %in% c( # Use Combine When the Additional Variable is Assoicated with Multiple Observations
"The Walt Disney Company",
"Comcast Corporation",
"Sony Group Corporation"
) ~ "Media & Entertainment",
parent %in% c(
"Volkswagen Group",
"General Motors"
) ~ "Automotive",
parent %in% c(
"Nestle",
"PepsiCo"
) ~ "Food & Beverage",
parent == "Johnson & Johnson" ~ "Healthcare", # Use Exactly Equal to When there is Only One Observation
parent == "JPMorgan Chase Co" ~ "Finance",
parent == "Birkshire Hathaway Inc" ~ "Conglomerate",
parent == "Walmart Inc" ~ "Retail",
# Set a Default Category for Uncategorized Brands
TRUE ~ "Other"
)
)
# Preview the New Data Frame with the Addition of the New Column
head(brands, 10)Data Note: Industry categories were assigned for this tutorial.
Most notably, the visNetwork package can be used to
customize the visual for meaningful analysis. The left_join
function is applied to add the newly created categorical variable of
“Industry” to the nodes. Use the left_join() function to
join the data set and the new data frame created above. In this example
the visNetwork and dplyr packages are employed
in tandem to customize and create more advanced interactive features.
This example demonstrates how visGroup(),
visOptions(), mutate() and
case_when() functions each contribute to interactive
visualization.
# Use the left_join() Function to Combine the New Data Frame with Original Data Set
nodes <- nodes %>%
left_join(
brands %>%
distinct(parent, industry),
by = c("id" = "parent")
) %>%
left_join(
brands %>%
select(brand, parent, industry) %>%
rename(brand.industry = industry),
by = c("id" = "brand")
) %>%
mutate(
industry = coalesce(industry, brand.industry),
# Customize the Size of the Nodes by Type
size = ifelse(
group == "Parent Company",
35,
15
),
# Add the Interactive Categorical Variables
title = ifelse(
group == "Parent Company",
paste0(
"<b>", label, "</b>",
"<br>Industry:", industry
),
paste0(
"<b>", label, "</b>",
"<br>Parent Company:", parent,
"<br>Industry:", industry
)
)
) %>%
select (-parent, -brand.industry)
#Use visNetwork() and visGroups() to Create the Interactive Network
visNetwork(nodes,
edges,
height = "650px",
width = "100%",
main = "Corportate Brand Relationships"
) %>%
visGroups( # Select the Variable Group to Customize
groupname = "Parent Company", # Name the Group/Selected Variables
color = "#177E89", # Set a Unique Color for this Group
shape = "dot" # Change the Shape for this Group
) %>%
visGroups(
groupname = "Brand",
color = "#E9B44C",
shape = "dot"
) %>%
# Use the visOptions() function to Add Interactive Selection
visOptions(
highlightNearest = TRUE,
nodesIdSelection = list(
enabled = TRUE,
main = "Select Brand Name:" # Name the Drop Down Box
),
selectedBy = list(
variable = "industry",
main = "Select an Industry:" # Name the Drop Down Box
)
) %>%
# Use visInteraction() to Add Hovering and Navigation
visInteraction(
hover = TRUE,
navigationButtons = TRUE
)
Learn more about the data set, how to create data frames, and the dplyr and visNetwork packahes and techniques with the following:
How to Use visNetwork Medium
visNetwork DataStorm/GitHub
stringsAsFactors blogr
Parent Companies and brands Data set
How to Create a Data Frame in R DataQuest
This code through references and cites the following sources:
Thieurmel, B. (2025). CRAN Introduction to visNetwork
Thieurmel, B., Arabiki, T., & Contributors (2022). GitHub. visNetwork
Beck, E.D. (n.d.). Data Manipulation: Intro to dplyr. Dplyr