Have you ever stared at your data and wondered what distribution it actually follows? Maybe it looks roughly normal, but the tail is too long. Maybe you have hydrology data and you know it’s skewed, but you’re not sure whether to go with a Gamma, a Log-Normal, or a Weibull. Or maybe you’ve been defaulting to the normal distribution simply because it’s the easiest option.
My colleague M. G. M. Khan and I have spent a lot of time in that situation, particularly when working with hydrological data, survey responses, and insurance claim sizes. That experience is what led us to build FitVerse, a new R package now available on CRAN.
The idea is simple: you hand FitVerse your data, and it tries a large number of distributions, estimates the parameters, ranks everything by how well it fits, and gives you a clean summary with plots and test statistics. No more fitting distributions one by one and manually comparing results.
FitVerse is on CRAN, so installation is straightforward:
install.packages("FitVerse")
library(FitVerse)
FitVerse covers 52 distribution families. That’s a wide range, and it’s intentional. Different fields tend to use different distributions, and we wanted one package to cover them all.
The distributions are organised into five groups:
Symmetric and unbounded: Normal, Logistic, Cauchy, Laplace, Student-t, Skew-Normal, Generalised Normal, Asymmetric Laplace, Johnson SU, Normal Inverse Gaussian, Generalised Hyperbolic, Truncated Normal.
Right-skewed with positive support: Gamma, Log-Normal, Exponential, Weibull (2P and 3P), Log-Logistic, Pareto, Inverse Gamma, Inverse Weibull, Paralogistic, Burr XII, Inverse Gaussian, Lomax, Generalised Gamma, Gompertz, Birnbaum-Saunders, Rayleigh, Half-Normal, Exponentiated Weibull, Gamma (3P), Log-Normal (3P), Lindley, Exponentiated Exponential, Nakagami, Dagum.
Heavy-tailed: GB2 (Generalised Beta of the Second Kind).
Extreme value and hydrological: GEV, Gumbel, GPD, Pearson Type III, Log-Pearson Type III, Generalised Logistic, Kappa-4, Wakeby.
Bounded and finite support: Beta, Uniform, Triangular, Right Triangular, PERT, Kumaraswamy.
FitVerse supports three ways of estimating parameters:
Maximum Likelihood (MLE) is the most widely used method and works well for most situations. It finds the parameter values that make your observed data most probable under the assumed distribution.
Method of Moments (MOM) matches theoretical moments (mean, variance) to the sample moments. It is simpler and often faster than MLE, and is a good choice when you want a quick, interpretable estimate.
L-Moments (LMOM) is especially popular in hydrology. L-moments are more robust to outliers than conventional moments, and they work well for fitting extreme-value distributions to things like flood peaks and rainfall maxima.
You can use any one of these, or run all three at once and let FitVerse pick the overall best fit across all methods and distributions.
Let’s use the built-in precip dataset, which has annual
precipitation values for 70 US cities. This is a right-skewed dataset,
so the normal distribution is unlikely to win.
x <- as.numeric(precip)
fit <- fitverse(x, xlab = "Annual precipitation (inches)")
That single line fits all 52 distributions using MLE, ranks them by AIC, runs goodness-of-fit tests, and produces a diagnostic plot. The printed output shows the top-ranked models, their parameters, AIC values, and test results.
To use all three methods:
fit <- fitverse(x, method = "ALL", xlab = "Annual precipitation (inches)")
If you already have candidate distributions in mind, you can pass a specific list:
fit <- fitverse(x,
dists = c("gamma", "lognormal", "weibull", "gev"),
method = "MLE")
To rank by BIC instead of AIC:
fit <- fitverse(x, criterion = "BIC")
Every call to fitverse() produces a two-panel diagnostic
plot by default.
The left panel is a histogram of your data with the top three fitted density curves overlaid. The best-fit distribution is highlighted, and its PDF formula is printed directly on the figure. This is useful for a quick visual check and for including in a report or paper.
The right panel is a Q-Q plot comparing the theoretical quantiles of the best-fit distribution to the sample quantiles. A good fit produces points that fall close to the diagonal line.
If you have plotly installed, you can get an interactive
version of the plot:
fit <- fitverse(x, interactive = TRUE)
This lets you zoom in, hover over points for values, and toggle distributions on and off.
To save the plot to a file without displaying it on screen:
save_plot(fit, file = "my_plot.png", dpi = 600)
You can also save as PDF or SVG by changing the file extension.
FitVerse runs four goodness-of-fit tests automatically: Kolmogorov-Smirnov, Anderson-Darling, Cramer-von Mises, and Chi-Squared. The results are shown in the printed summary and included in the ranking table.
The composite GoF score is used alongside AIC and BIC to help you make a more informed decision. A distribution might have the lowest AIC but still fail the Anderson-Darling test, and FitVerse makes that visible rather than hiding it.
Once you have a fit, you can compute bootstrap confidence intervals for the estimated parameters:
boot <- bootstrap_ci(fit, R = 500)
print(boot)
You can also compute return levels, which are useful in hydrology and risk analysis. A T-year return level is the value expected to be exceeded once every T years on average:
rl <- return_level(fit, T = c(10, 50, 100))
print(rl)
Bootstrap CIs for return levels are also supported:
boot_rl <- bootstrap_ci(fit, R = 500, return_levels = c(10, 50, 100))
print(boot_rl)
After fitting, you can use the best-fit distribution directly for probability calculations:
# Density at x = 30
dfitverse(30, fit)
# Probability that x is less than 40
pfitverse(40, fit)
# 95th percentile
qfitverse(0.95, fit)
# Generate 100 random values from the fitted distribution
rfitverse(100, fit)
These work for all 52 distributions without you needing to know the underlying function names.
If you are fitting extreme-value or hydrological distributions, the L-Moment Ratio Diagram (LMRD) is a useful exploratory tool. It plots your sample L-skewness and L-kurtosis against the theoretical curves for several distribution families, giving you a visual guide to which family is likely to fit best before you even run the fitting.
lmrd(x)
If you have a data frame with multiple numeric columns and want to
fit distributions to all of them at once, fitverse_batch()
handles this:
results <- fitverse_batch(my_data, method = "MLE")
results$summary
The summary table gives you the best-fit distribution and parameters for each column side by side. This is handy for datasets where you have many variables and want a quick overview of their distributional shapes.
FitVerse can produce a self-contained HTML report of your analysis. It includes the fitted parameters, rankings, test results, and the diagnostic plot, all in one file you can share or archive.
generate_report(fit,
title = "Precipitation Data: Distribution Fitting",
author = "Karuna Reddy",
output_file = file.path(tempdir(), "report.html"))
If you have rmarkdown installed it renders a polished
HTML document. If not, it produces a plain HTML fallback. You can also
generate a PDF by changing the file extension to .pdf.
To include bootstrap results in the report:
boot <- bootstrap_ci(fit, R = 500)
generate_report(fit, boot = boot, output_file = file.path(tempdir(), "report.html"))
For users who prefer not to write code, FitVerse ships with a full Shiny application. You launch it with:
FitVerseApp()
The app lets you upload a CSV file, choose which distributions to fit, switch between MLE, MOM, and LMOM, explore the diagnostic plots interactively, and download the results as a report. It is designed for teaching and for sharing with colleagues who do not use R directly.
FitVerse is designed for anyone who works with continuous univariate data and needs to characterise its distribution. Some specific use cases:
Hydrology: fitting extreme-value distributions (GEV, Gumbel, GPD, Pearson III) to flood peak data, rainfall maxima, or drought indices. The LMOM method and return level functions make FitVerse a natural tool for frequency analysis.
Actuarial science: fitting heavy-tailed distributions (Pareto, Burr, GB2, Log-Normal) to insurance claim sizes or loss data.
Survey sampling: identifying the distribution of a continuous survey variable before applying model-based estimation or simulation.
Teaching: the app and the clean printed output make FitVerse useful in a classroom setting, where students can explore distribution fitting visually without needing to know the underlying code.
Research: the combination of 52 distributions, three estimation methods, four GoF tests, bootstrap CIs, and report generation covers most of what a researcher needs when reporting a distributional analysis in a paper.
If you use FitVerse in your work, please cite it. You can get the citation in R with:
citation("FitVerse")
FitVerse is on CRAN: https://cran.r-project.org/package=FitVerse
We would love to hear how you are using it. Feel free to leave a comment below or get in touch directly.