Hypergeometric Distribution

Joel Sanqui (Appalachian State University)

Introduction

Let (X) be a random variable representing the total number of successes in a sample drawn without replacement from a finite population.

  • The total population size is (N).
  • The total number of success states in the population is (M).
  • The total number of failure states in the population is (N - M).
  • We draw a sample of size (n) from this population.

Finding the pmf

To find the probability of getting exactly (x) successes in our sample, we use combinatorics to count the favorable outcomes over the total possible outcomes:

\[P(X = x)\] \[=\frac{\mathbf{\text{Ways to choose successes}} \times \mathbf{\text{Ways to choose failures}}}{\mathbf{\text{Total ways to choose the sample}}}\]

1. Counting the Combinations

We must choose exactly \(x\) successes from the \(M\) available, and the remaining \(n-x\) items must be failures chosen from the \(N-M\) available.

  • Ways to choose successes: \(\binom{M}{x}\)
  • Ways to choose failures: \(\binom{N-M}{n-x}\)
  • Total ways to choose any sample of size \(n\): \(\binom{N}{n}\)

2. The Final pmf

Combining these combinations gives the probability mass function for the hypergeometric distribution:

\[f(x) = P(X = x) = \frac{\binom{M}{x} \binom{N-M}{n-x}}{\binom{N}{n}}\]

\[\text{for } \max(0, n - (N - M)) \le x \le \min(n, M)\]

We write \[X \sim \text{Hypergeometric}(N, M, n)\].

Mean of Hypergeometric

Let \[p = \frac{M}{N}\] be the initial probability of success. Even though the trials are dependent, by the linearity of expectation, the expected value matches the binomial form:

\[\mathbb{E}[X] = n \cdot \frac{M}{N} = np\]

Variance of Hypergeometric

Because drawing is done without replacement, the trials are dependent. This introduces a finite population correction (FPC) factor, \[\frac{N-n}{N-1}\], which reduces the variance compared to the binomial distribution:

\[\text{Var}(X) = n \left(\frac{M}{N}\right) \left(1 - \frac{M}{N}\right) \left(\frac{N-n}{N-1}\right)\]

Example 1: Quality Control

Manufacturing plants use the hypergeometric distribution when sampling without replacement from a small, finite batch.

  • Scenario: A factory inspects a specific batch of microchips.
  • Parameters: A batch contains \(N = 50\) total microchips, where exactly \(M = 5\) are defective. A quality engineer samples \(n = 10\) microchips without replacement.
  • Problem: Calculate the probability that \(x \le 2\) chips are defective, and find the mean and variance of the defective chips in the sample.

Solutions

  • Exact Probability: \(P(X \le 2) = f(0)+f(1)+f(2)\) = 95.17%
  • Mean: Expected number of defective chips is \(10 \cdot \left(\frac{5}{50}\right) =\) 1.00
  • Variance: \[10 \cdot \left(\frac{5}{50}\right) \cdot \left(\frac{45}{50}\right) \cdot \left(\frac{50-10}{50-1}\right) \approx\] 0.7347

QC Visualization

library(ggplot2)

N <- 50   # Total population
M <- 5    # Successes in population (defects)
n_failures <- N - M  # Failures in population
n <- 10   # Sample size

df_qc <- data.frame(x = 0:5, pmf = dhyper(0:5, m = M, n = n_failures, k = n))
df_qc$highlight <- ifelse(df_qc$x <= 2, "Target (<=2)", "Other")

ggplot(df_qc, aes(x = factor(x), y = pmf, fill = highlight)) +
  geom_col(width = 0.7) +
  scale_fill_manual(values = c("Target (<=2)" = "darkgreen", "Other" = "grey70")) +
  labs(title = "QC Hypergeometric Probability Matrix (N=50, M=5, n=10)",
       x = "Number of Defective Chips (x)", y = "Probability", fill = "Region") +
  theme_minimal()

QC Visualization

Example 2: Committee Selection

Auditors and researchers use this distribution when sampling from human populations where individual status cannot be replicated.

  • Scenario: An organization selects a small committee from a local pool of registered voters.
  • Parameters: In a small district of \(N = 500\) voters, \(M = 200\) support Candidate A. A random panel of \(n = 50\) voters is chosen without replacement.
  • Question: How likely is it to observe \(x \ge 25\) supporters of Candidate A in the panel?

Solution

  • Exact Hypergeometric Probability: \(P(X \ge 25) = 1 - P(X \le 24)\) = 8.63%
  • Mean: Expected number of supporters is \(50 \cdot \left(\frac{200}{500}\right) =\) 20
  • Variance: \(50 \cdot (0.4) \cdot (0.6) \cdot \left(\frac{500-50}{500-1}\right) \approx\) 10.82

Panel Selection Visualization

library(ggplot2)

N <- 500
M <- 200
n_failures <- N - M
n <- 50

df_panel <- data.frame(x = 5:35, pmf = dhyper(5:35, m = M, n = n_failures, k = n))
df_panel$highlight <- ifelse(df_panel$x >= 25, "Target (>=25)", "Other")

ggplot(df_panel, aes(x = x, y = pmf, fill = highlight)) +
  geom_col(width = 0.9) +
  scale_fill_manual(values = c("Target (>=25)" = "firebrick", "Other" = "grey70")) +
  labs(title = "Committee Selection (N=500, M=200, n=50)",
       x = "Number of Supporters (x)", y = "Probability", fill = "Region") +
  theme_minimal()

Panel Selection Visualization

Hypergeometric Distribution in R

In R, the arguments for hypergeometric functions are structured as follows: - m: Number of success states in the population (\(M\)) - n: Number of failure states in the population (\(N - M\)) - k: Sample size (\(n\)))

The probability from Example 1 using dhyper is:

dhyper(0, m=5, n=45, k=10) + dhyper(1, m=5, n=45, k=10) + dhyper(2, m=5, n=45, k=10)
[1] 0.9517397

Using the cumulative distribution function phyper:

phyper(2, m=5, n=45, k=10)
[1] 0.9517397

Quantiles and random numbers are generated using qhyper and rhyper.

Example 3

What is the probability of drawing exactly 5 red cards if 10 cards are dealt from a standard, well-shuffled 52-card deck without replacement?

Solution:

Let \(X\) = number of red cards. Population \(N=52\), Successes \(M=26\) (red cards), Failures \(N-M=26\) (black cards), Sample size \(n=10\).

dhyper(5, m = 26, n = 26, k = 10)
[1] 0.2735147

Example 4

What is the probability of getting less than 4 red cards?

Solution:

\(P(X < 4) = P(X \le 3)\) is calculated as:

phyper(3, m = 26, n = 26, k = 10)
[1] 0.1456742

Example 5

What is the probability of getting anywhere from 4 to 6 red cards?

Solution:

\(P(4 \le X \le 6)\) is calculated as:

phyper(6, m = 26, n = 26, k = 10) - phyper(3, m = 26, n = 26, k = 10)
[1] 0.7086516

Example 6

What is the 3rd quartile of this card-drawing distribution?

Solution:

The 75th quantile is calculated as:

qhyper(0.75, m = 26, n = 26, k = 10)
[1] 6

Empirical Simulation: Card Dealing

What happens if we simulate a Hypergeometric process empirically?

  • Experiment: Deal \(10\) cards from a standard deck (\(M=26\) red, \(N-M=26\) black).
  • Repetitions: Repeat this dealing process \(1,000\) times.
  • Objective: Approximate the distribution of the random variable that counts the number of red cards dealt and estimate its mean and variance from the simulation.

Simulation & Analysis

library(ggplot2)
set.seed(42) # Ensure reproducible results

# 1000 repetitions, drawing 10 cards from a deck of 26 red and 26 black
sim_deals <- rhyper(nn = 1000, m = 26, n = 26, k = 10)
df_sim <- data.frame(red_cards = sim_deals)

# Calculate empirical values
sim_mean <- mean(sim_deals)
sim_var  <- var(sim_deals)

# Dynamic Title with results
plot_title <- paste0("Empirical Card Deals (1000 trials of n=10)\n",
                     "Simulated Mean: ", round(sim_mean, 2), 
                     " | Simulated Variance: ", round(sim_var, 2))

# Plot histogram
ggplot(df_sim, aes(x = red_cards)) +
  geom_histogram(binwidth = 1, fill = "purple", color = "white", alpha = 0.8) +
  geom_vline(xintercept = sim_mean, color = "black", linetype = "dashed", size = 1) +
  labs(title = plot_title, x = "Number of Red Cards (X)", y = "Frequency") +
  theme_minimal()

Simulation & Analysis

Note that the theoretical mean is 5 and the theoretical variance is \(\approx 2.196\) (\(10 \cdot 0.5 \cdot 0.5 \cdot \frac{42}{51}\)), so the simulation is pretty good.

References

  • History of the Distribution: Hald, A. (1998). A History of Mathematical Statistics from 1750 to 1930. New York, NY: John Wiley & Sons.

  • Properties of the Distribution Ross, S. M. (2014). A First Course in Probability (9th ed.). Pearson.

  • Computational Software: R Core Team (2026). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. Vienna, Austria. Available at https://www.R-project.org/.