Let (X) be a random variable representing the total number of successes in a sample drawn without replacement from a finite population.
The total population size is (N).
The total number of success states in the population is (M).
The total number of failure states in the population is (N - M).
We draw a sample of size (n) from this population.
Finding the pmf
To find the probability of getting exactly (x) successes in our sample, we use combinatorics to count the favorable outcomes over the total possible outcomes:
\[P(X = x)\]\[=\frac{\mathbf{\text{Ways to choose successes}} \times \mathbf{\text{Ways to choose failures}}}{\mathbf{\text{Total ways to choose the sample}}}\]
1. Counting the Combinations
We must choose exactly \(x\) successes from the \(M\) available, and the remaining \(n-x\) items must be failures chosen from the \(N-M\) available.
Ways to choose successes: \(\binom{M}{x}\)
Ways to choose failures: \(\binom{N-M}{n-x}\)
Total ways to choose any sample of size \(n\): \(\binom{N}{n}\)
2. The Final pmf
Combining these combinations gives the probability mass function for the hypergeometric distribution:
\[\text{for } \max(0, n - (N - M)) \le x \le \min(n, M)\]
We write \[X \sim \text{Hypergeometric}(N, M, n)\].
Mean of Hypergeometric
Let \[p = \frac{M}{N}\] be the initial probability of success. Even though the trials are dependent, by the linearity of expectation, the expected value matches the binomial form:
\[\mathbb{E}[X] = n \cdot \frac{M}{N} = np\]
Variance of Hypergeometric
Because drawing is done without replacement, the trials are dependent. This introduces a finite population correction (FPC) factor, \[\frac{N-n}{N-1}\], which reduces the variance compared to the binomial distribution:
\[\text{Var}(X) = n \left(\frac{M}{N}\right) \left(1 - \frac{M}{N}\right) \left(\frac{N-n}{N-1}\right)\]
Example 1: Quality Control
Manufacturing plants use the hypergeometric distribution when sampling without replacement from a small, finite batch.
Scenario: A factory inspects a specific batch of microchips.
Parameters: A batch contains \(N = 50\) total microchips, where exactly \(M = 5\) are defective. A quality engineer samples \(n = 10\) microchips without replacement.
Problem: Calculate the probability that \(x \le 2\) chips are defective, and find the mean and variance of the defective chips in the sample.
library(ggplot2)N <-50# Total populationM <-5# Successes in population (defects)n_failures <- N - M # Failures in populationn <-10# Sample sizedf_qc <-data.frame(x =0:5, pmf =dhyper(0:5, m = M, n = n_failures, k = n))df_qc$highlight <-ifelse(df_qc$x <=2, "Target (<=2)", "Other")ggplot(df_qc, aes(x =factor(x), y = pmf, fill = highlight)) +geom_col(width =0.7) +scale_fill_manual(values =c("Target (<=2)"="darkgreen", "Other"="grey70")) +labs(title ="QC Hypergeometric Probability Matrix (N=50, M=5, n=10)",x ="Number of Defective Chips (x)", y ="Probability", fill ="Region") +theme_minimal()
QC Visualization
Example 2: Committee Selection
Auditors and researchers use this distribution when sampling from human populations where individual status cannot be replicated.
Scenario: An organization selects a small committee from a local pool of registered voters.
Parameters: In a small district of \(N = 500\) voters, \(M = 200\) support Candidate A. A random panel of \(n = 50\) voters is chosen without replacement.
Question: How likely is it to observe \(x \ge 25\) supporters of Candidate A in the panel?
library(ggplot2)N <-500M <-200n_failures <- N - Mn <-50df_panel <-data.frame(x =5:35, pmf =dhyper(5:35, m = M, n = n_failures, k = n))df_panel$highlight <-ifelse(df_panel$x >=25, "Target (>=25)", "Other")ggplot(df_panel, aes(x = x, y = pmf, fill = highlight)) +geom_col(width =0.9) +scale_fill_manual(values =c("Target (>=25)"="firebrick", "Other"="grey70")) +labs(title ="Committee Selection (N=500, M=200, n=50)",x ="Number of Supporters (x)", y ="Probability", fill ="Region") +theme_minimal()
Panel Selection Visualization
Hypergeometric Distribution in R
In R, the arguments for hypergeometric functions are structured as follows: - m: Number of success states in the population (\(M\)) - n: Number of failure states in the population (\(N - M\)) - k: Sample size (\(n\)))
Using the cumulative distribution function phyper:
phyper(2, m=5, n=45, k=10)
[1] 0.9517397
Quantiles and random numbers are generated using qhyper and rhyper.
Example 3
What is the probability of drawing exactly 5 red cards if 10 cards are dealt from a standard, well-shuffled 52-card deck without replacement?
Solution:
Let \(X\) = number of red cards. Population \(N=52\), Successes \(M=26\) (red cards), Failures \(N-M=26\) (black cards), Sample size \(n=10\).
dhyper(5, m =26, n =26, k =10)
[1] 0.2735147
Example 4
What is the probability of getting less than 4 red cards?
Solution:
\(P(X < 4) = P(X \le 3)\) is calculated as:
phyper(3, m =26, n =26, k =10)
[1] 0.1456742
Example 5
What is the probability of getting anywhere from 4 to 6 red cards?
Solution:
\(P(4 \le X \le 6)\) is calculated as:
phyper(6, m =26, n =26, k =10) -phyper(3, m =26, n =26, k =10)
[1] 0.7086516
Example 6
What is the 3rd quartile of this card-drawing distribution?
Solution:
The 75th quantile is calculated as:
qhyper(0.75, m =26, n =26, k =10)
[1] 6
Empirical Simulation: Card Dealing
What happens if we simulate a Hypergeometric process empirically?
Experiment: Deal \(10\) cards from a standard deck (\(M=26\) red, \(N-M=26\) black).
Repetitions: Repeat this dealing process \(1,000\) times.
Objective: Approximate the distribution of the random variable that counts the number of red cards dealt and estimate its mean and variance from the simulation.
Simulation & Analysis
library(ggplot2)set.seed(42) # Ensure reproducible results# 1000 repetitions, drawing 10 cards from a deck of 26 red and 26 blacksim_deals <-rhyper(nn =1000, m =26, n =26, k =10)df_sim <-data.frame(red_cards = sim_deals)# Calculate empirical valuessim_mean <-mean(sim_deals)sim_var <-var(sim_deals)# Dynamic Title with resultsplot_title <-paste0("Empirical Card Deals (1000 trials of n=10)\n","Simulated Mean: ", round(sim_mean, 2), " | Simulated Variance: ", round(sim_var, 2))# Plot histogramggplot(df_sim, aes(x = red_cards)) +geom_histogram(binwidth =1, fill ="purple", color ="white", alpha =0.8) +geom_vline(xintercept = sim_mean, color ="black", linetype ="dashed", size =1) +labs(title = plot_title, x ="Number of Red Cards (X)", y ="Frequency") +theme_minimal()
Simulation & Analysis
Note that the theoretical mean is 5 and the theoretical variance is \(\approx 2.196\) (\(10 \cdot 0.5 \cdot 0.5 \cdot \frac{42}{51}\)), so the simulation is pretty good.
References
History of the Distribution: Hald, A. (1998). A History of Mathematical Statistics from 1750 to 1930. New York, NY: John Wiley & Sons.
Properties of the Distribution Ross, S. M. (2014). A First Course in Probability (9th ed.). Pearson.
Computational Software: R Core Team (2026). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. Vienna, Austria. Available at https://www.R-project.org/.