This report explores the FIFA 18 Complete Player
Dataset, published on Kaggle and originally scraped from sofifa.com, the statistics database behind
EA Sports’ FIFA video game series. The file used here,
CompleteDataset.csv, contains 17,981 professional
and semi-professional football (soccer) players rated by EA
Sports for the FIFA 18 edition of the game, along with 73 additional
variables describing each player.
The variables fall into a few broad groups:
Name,
Age, Nationality, Club.Overall (current
ability, 0-99), Potential (ceiling ability), and
Special (a composite index).Value and
Wage, stored as text such as "€95.5M" or
"€565K".Acceleration, Dribbling,
Finishing, Strength, and Vision,
each on a 0-99 scale.ST,
CM, CB), plus a
Preferred Positions column listing the position(s) a player
is best suited to.I chose this dataset because it combines simple demographic variables (age, nationality) with a rich set of numeric performance ratings, which makes it well suited to demonstrating several different kinds of descriptive visualization – distributions, group comparisons, relationships between variables, categorical counts, and correlation structure.
No personally sensitive information is involved: every player in the dataset is a public professional athlete, and all fields describe publicly published, game-related ratings rather than private information.
Three cleaning steps were necessary before any analysis could be done. The code below performs all of them; click “Code” to expand it if you would like to see exactly how the data was prepared.
Value and Wage out
of text like "€105M" / "€565K" into plain
numeric Euro amounts."80+3" – a base rating plus a positional adjustment EA
Sports applies for out-of-position players. Converting these directly
with as.numeric() turns just those specific values into
NA, which R’s summary functions and ggplot2
already know how to skip over."CB", "ST") is mapped
to one of four broad roles: Goalkeeper, Defender, Midfielder, or
Forward.fifa_raw <- read_csv("CompleteDataset.csv", show_col_types = FALSE)
# Step 1: drop the unnamed index column
fifa_raw <- fifa_raw %>% select(-1)
# Make column names easier to reference (spaces -> underscores)
names(fifa_raw) <- str_replace_all(names(fifa_raw), " ", "_")
# Step 2: parse Value / Wage text into numeric Euros.
# parse_number() (from readr) reads off the leading number and ignores
# the currency symbol and any letters, so "€95.5M" becomes 95.5 and
# "€0" becomes 0. We then multiply by the right amount depending on
# whether the original text ended in "M" (millions) or "K" (thousands).
parse_money <- function(x) {
amount <- parse_number(x)
multiplier <- case_when(
str_ends(x, "M") ~ 1e6,
str_ends(x, "K") ~ 1e3,
TRUE ~ 1
)
amount * multiplier
}
# Which columns hold the 0-99 skill ratings (Acceleration, Dribbling, etc.)
skill_cols <- c("Acceleration", "Aggression", "Agility", "Balance",
"Ball_control", "Composure", "Crossing", "Curve",
"Dribbling", "Finishing", "Free_kick_accuracy",
"Heading_accuracy", "Interceptions", "Jumping",
"Long_passing", "Long_shots", "Marking", "Penalties",
"Positioning", "Reactions", "Short_passing",
"Shot_power", "Sliding_tackle", "Sprint_speed",
"Stamina", "Standing_tackle", "Strength", "Vision",
"Volleys")
fifa <- fifa_raw %>%
mutate(
Value_EUR = parse_money(Value),
Wage_EUR = parse_money(Wage),
Value_M = Value_EUR / 1e6,
Wage_K = Wage_EUR / 1e3,
First_Position = word(str_trim(Preferred_Positions), 1)
) %>%
mutate(across(all_of(skill_cols), as.numeric)) %>%
# Step 4: map each player's first preferred position to a broad role
mutate(Position_Category = case_when(
First_Position == "GK" ~ "Goalkeeper",
First_Position %in% c("CB", "LB", "RB", "LWB", "RWB", "LCB", "RCB") ~ "Defender",
First_Position %in% c("CDM", "CM", "CAM", "LM", "RM", "LAM", "RAM",
"LDM", "RDM", "LCM", "RCM") ~ "Midfielder",
First_Position %in% c("ST", "CF", "LW", "RW", "LF", "RF", "LS", "RS") ~ "Forward",
TRUE ~ NA_character_
)) %>%
mutate(Position_Category = factor(
Position_Category,
levels = c("Goalkeeper", "Defender", "Midfielder", "Forward")
))
After cleaning, the working dataset has 17981 players and 80 columns. Here is a preview of the key columns used in the rest of this report:
fifa %>%
select(Name, Age, Nationality, Club, Overall, Potential,
Value_M, Position_Category) %>%
head(10)
Before building any visualizations, it is useful to look at the basic
numeric summaries of the variables at the center of this analysis:
Age, Overall rating, Potential,
and Value_M (market value in millions of Euros).
summary(fifa %>% select(Age, Overall, Potential, Value_M))
## Age Overall Potential Value_M
## Min. :16.00 Min. :46.00 Min. :46.00 Min. : 0.000
## 1st Qu.:21.00 1st Qu.:62.00 1st Qu.:67.00 1st Qu.: 0.300
## Median :25.00 Median :66.00 Median :71.00 Median : 0.675
## Mean :25.14 Mean :66.25 Mean :71.19 Mean : 2.385
## 3rd Qu.:28.00 3rd Qu.:71.00 3rd Qu.:75.00 3rd Qu.: 2.100
## Max. :47.00 Max. :94.00 Max. :94.00 Max. :123.000
fifa %>%
summarise(
n = n(),
mean_age = round(mean(Age), 1),
sd_age = round(sd(Age), 1),
mean_overall = round(mean(Overall), 1),
median_overall = median(Overall),
sd_overall = round(sd(Overall), 1),
mean_value_M = round(mean(Value_M), 2),
median_value_M = round(median(Value_M), 2),
max_value_M = round(max(Value_M), 1)
)
A few points stand out immediately:
Overall rating is almost perfectly symmetric: the mean
(66.2) and median (66) are nearly identical, which we will see confirmed
in the histogram below.Value_M is extremely right-skewed: the median player is
worth well under €1 million, while the most valuable player in the
dataset is worth €123 million. A handful of superstars
pull the mean far above the median.fifa %>% count(Position_Category, sort = TRUE)
fifa %>% count(Nationality, sort = TRUE) %>% slice_head(n = 10)
With these baseline numbers established, the rest of the report walks through five different visualizations, each chosen to highlight a different aspect of the data: an overall distribution, a comparison across groups, a relationship between two numeric variables, a ranked categorical count, and a correlation structure among many variables at once.
p1 <- ggplot(fifa, aes(x = Overall)) +
geom_histogram(binwidth = 2, fill = "#2C5F8A", color = "white") +
geom_vline(aes(xintercept = mean(Overall)),
color = "#D9534F", linetype = "dashed", linewidth = 1) +
geom_vline(aes(xintercept = median(Overall)),
color = "#F0AD4E", linetype = "solid", linewidth = 1) +
labs(
title = "Distribution of Player Overall Ratings",
subtitle = "Dashed red line = mean, solid orange line = median",
x = "Overall Rating", y = "Number of Players"
) +
theme_minimal(base_size = 13)
p1
Histogram of player Overall ratings
What this shows: A histogram is the natural first
visualization for any numeric variable because it shows the
shape of the distribution directly – how values are spread out,
whether the distribution is symmetric or skewed, and whether there are
multiple peaks. Here, each bar groups players into a 2-point-wide band
of Overall rating and shows how many players fall into that
band.
Design choices: A bin width of 2 was chosen because
Overall is an integer rating from roughly 46 to 94; wider
bins would hide the shape of the distribution, while narrower bins (bin
width of 1) produced a noisy, “toothy” histogram without adding useful
detail. The dashed red line marks the mean and the solid orange line
marks the median so the reader can judge symmetry at a glance.
Interpretation: The distribution is unimodal and close to symmetric – the mean and median lines sit almost exactly on top of each other around a rating of 66. This makes sense given how EA Sports designs its rating system: most professional players cluster in the 60s, with a long-ish right tail representing the small number of elite, world-class players (ratings in the high 80s and 90s) and a shorter left tail representing fringe or lower-league players. There is no evidence of multiple peaks, which would have suggested two distinct populations of players (for example, if top-league and lower-league players had been rated on very different scales).
p2 <- ggplot(fifa, aes(x = fct_reorder(Position_Category, Value_M, .fun = median),
y = Value_M, fill = Position_Category)) +
geom_boxplot(outlier.alpha = 0.3, show.legend = FALSE) +
scale_y_log10(labels = label_number(suffix = "M")) +
scale_fill_brewer(palette = "Blues") +
labs(
title = "Market Value by Position Category",
subtitle = "Log scale used because a small number of superstar players\nare worth far more than the typical player",
x = NULL, y = "Market Value (millions of Euros, log scale)"
) +
theme_minimal(base_size = 13)
p2
Boxplot of market value by position category
What this shows: A boxplot summarizes the distribution of a numeric variable (market value) separately for each level of a categorical variable (position category), making it easy to compare center, spread, and outliers across groups side by side. Each box spans the interquartile range (25th to 75th percentile), the line inside the box is the median, the whiskers extend to roughly 1.5 times the interquartile range, and individual points beyond the whiskers are flagged as outliers.
Design choices: Market value is extremely right-skewed (as noted in the descriptive statistics above), so the y-axis is shown on a log scale. Without it, the handful of €100M+ superstars would compress every other box into an unreadable sliver near zero. The four position categories are ordered from lowest to highest median value (left to right) so the reader can immediately see the ranking, and each box is filled with a shade from the same color family to keep the chart visually calm rather than using four unrelated colors.
Interpretation: Goalkeepers have the lowest typical market value, consistent with the tendency of a single goalkeeper to command a smaller transfer fee than an outfield attacking player. Forwards and midfielders have the highest median values and also the widest spread and the most extreme high-end outliers – reflecting how the game’s (and football’s) biggest transfer fees and highest wages tend to go to attacking players who can single-handedly change the outcome of a match. Every position group still shows substantial overlap and a long tail of high-value outliers, which is expected: even among defenders, a small number of world-class players are worth far more than a typical squad player at the same position.
p3 <- ggplot(fifa, aes(x = Age, y = Overall)) +
geom_point(alpha = 0.12, size = 0.8, color = "#2C5F8A") +
geom_smooth(method = "loess", color = "#D9534F", se = FALSE, linewidth = 1.1) +
facet_wrap(~ Position_Category) +
labs(
title = "Age vs. Overall Rating by Position",
subtitle = "Smoothed trend line shows the typical career arc",
x = "Age", y = "Overall Rating"
) +
theme_minimal(base_size = 13)
p3
Scatter plot of Age vs Overall rating, faceted by position
What this shows: A scatter plot displays the relationship between two continuous numeric variables – here, a player’s age and their overall rating – with each point representing one player. Faceting (splitting the plot into small panels, one per position category) lets us check whether that relationship looks the same for goalkeepers, defenders, midfielders, and forwards, or whether it differs by role.
Design choices: With almost 18,000 players, plotting
every point at full opacity would produce a solid, uninformative blob.
Setting alpha (transparency) to about 0.12 lets overlapping
points show through as darker regions, effectively turning the scatter
plot into an informal density map while still preserving individual
points. A LOESS-smoothed trend line (shown in red) is layered on top of
each panel to summarize the central tendency of the relationship without
assuming it has to be a straight line.
Interpretation: All four panels show the same
characteristic inverted-U (“arc”) shape: overall rating rises quickly
through a player’s late teens and early twenties, peaks somewhere around
age 29-31, and then gradually declines. This matches real-world
intuition about athletic careers – players improve as they gain
experience and physical maturity, plateau during their prime years, and
decline as age affects speed, recovery, and physical conditioning. The
correlation between Age and Overall across the
whole dataset is 0.46 (moderate and positive), but the
scatter plot reveals why a single correlation number is
misleading here: the true relationship is curved, not a straight line,
so a linear correlation coefficient only partially captures the pattern.
The shape is strikingly consistent across all four position groups,
suggesting the aging curve is a general athletic phenomenon rather than
something specific to one style of play.
top_nat <- fifa %>%
count(Nationality, sort = TRUE) %>%
slice_head(n = 15)
p4 <- ggplot(top_nat, aes(x = fct_reorder(Nationality, n), y = n)) +
geom_col(fill = "#2C5F8A") +
coord_flip() +
labs(
title = "Top 15 Nationalities by Player Count",
x = NULL, y = "Number of Players"
) +
theme_minimal(base_size = 13)
p4
Bar chart of the top 15 nationalities by player count
What this shows: A bar chart is the standard tool for comparing counts (or totals) across categories. Here, each bar represents one nationality, and its length shows how many players of that nationality appear in the dataset.
Design choices: With over 150 distinct nationalities
in the full dataset, showing all of them would be unreadable, so the
chart is limited to the top 15 by count. The bars are horizontal
(coord_flip()) so the (sometimes long) country names remain
readable without rotating text, and they are sorted from least to
greatest so the largest bar naturally appears at the top of the chart –
a small design choice that makes the ranking much easier to read at a
glance than an alphabetical ordering would.
Interpretation: England is by far the most represented nationality, followed by Germany, Spain, France, and Argentina. This does not necessarily mean England produces more elite footballers than other countries in an absolute sense – it more likely reflects coverage bias in the underlying data source: sofifa.com (and the FIFA game itself) is especially thorough in cataloguing English lower-league clubs (League One, League Two, and non-league sides) in addition to the Premier League, which inflates England’s total player count relative to countries where only the top one or two divisions are represented in the game. This is a useful reminder that a simple count-based bar chart can reflect quirks of how data was collected just as much as it reflects the real-world phenomenon being measured.
corr_attrs <- c("Acceleration", "Agility", "Balance", "Dribbling",
"Finishing", "Short_passing", "Strength", "Stamina",
"Aggression", "Vision")
corr_matrix <- fifa %>%
select(all_of(corr_attrs)) %>%
cor(use = "pairwise.complete.obs")
# Turn the correlation matrix (a grid) into one row per pair of
# attributes, which is the shape ggplot2 needs for geom_tile().
# as.table() + as.data.frame() does this in one step, giving us
# columns Var1, Var2, and Freq (the correlation value).
corr_long <- as.data.frame(as.table(corr_matrix))
names(corr_long) <- c("Attribute_1", "Attribute_2", "Correlation")
corr_long$Attribute_1 <- factor(corr_long$Attribute_1, levels = corr_attrs)
corr_long$Attribute_2 <- factor(corr_long$Attribute_2, levels = corr_attrs)
p5 <- ggplot(corr_long, aes(x = Attribute_2, y = Attribute_1, fill = Correlation)) +
geom_tile(color = "white") +
geom_text(aes(label = sprintf("%.2f", Correlation)), size = 3) +
scale_fill_gradient2(low = "#2166AC", mid = "white", high = "#B2182B",
midpoint = 0, limits = c(-1, 1)) +
labs(title = "Correlation Among Skill Attributes", x = NULL, y = NULL) +
theme_minimal(base_size = 12) +
theme(axis.text.x = element_text(angle = 45, hjust = 1))
p5
Correlation heatmap among ten skill attributes
What this shows: A correlation heatmap compresses an entire correlation matrix – every pairwise correlation among a set of numeric variables – into a single grid, using color (and, here, printed numbers) to encode the strength and direction of each relationship. This makes it possible to spot patterns across many variables at once that would be tedious to find by scanning individual scatter plots one pair at a time.
Design choices: Ten representative skill attributes
were selected out of the 29 available, chosen to span physical
attributes (Acceleration, Agility,
Balance, Strength, Stamina),
technical attributes (Dribbling, Finishing,
Short_passing), and mental/behavioral attributes
(Aggression, Vision) – using all 29 would have
produced an overwhelming 29x29 grid. A diverging color scale (blue for
negative correlation, white for none, red for positive) is used with the
midpoint fixed at zero, and the actual correlation coefficient is
printed inside each tile so exact values are readable rather than only
approximated by color.
Interpretation: Most technical and physical
attributes are positively correlated with one another – for example,
Dribbling, Finishing,
Short_passing, and Vision form a tight cluster
of attributes (pairwise correlations mostly in the 0.65-0.85 range),
reflecting the fact that technically gifted attacking players tend to be
rated highly across all of these skills together. Strength,
by contrast, stands out as the one attribute with
negative correlations against Agility,
Balance, and Acceleration (roughly -0.16 to
-0.40) – a sensible finding, since bulkier, more physically powerful
players tend to be comparatively less quick and less balanced than
smaller, more agile players. Aggression correlates most
strongly with Stamina and Strength rather than
with the technical cluster, consistent with aggression being more
closely associated with physical, combative play styles (common among
defenders and defensive midfielders) than with finesse-based attacking
skills.
This report examined the FIFA 18 Complete Player Dataset using five different types of visualization, each suited to a different descriptive question: a histogram to reveal the shape of a single distribution, a boxplot to compare a numeric variable across groups, a scatter plot (with faceting and a trend line) to explore a relationship between two numeric variables, a bar chart to rank categorical counts, and a correlation heatmap to summarize relationships among many variables simultaneously.
Together, these visualizations tell a coherent story about the dataset: player ability ratings are approximately normally distributed around a league-average rating in the mid-60s; market value is highly concentrated among a small number of elite (often attacking) players; player ability follows a predictable age-related arc that peaks around age 30 regardless of position; the dataset’s national coverage is uneven in a way that reflects data-collection choices rather than football talent alone; and skill attributes cluster into recognizable physical, technical, and behavioral groups, with a few attributes (like strength versus agility) trading off against one another.
Data source: FIFA 18 Complete Player Dataset, Kaggle, originally scraped from sofifa.com.