3.1 3.1 Primary Objective
To estimate the association/effect of [primary exposure] on [primary outcome] in [target population].
Study title: [Study Title]
Protocol version/date: [Protocol Version]
SAP version: 1.0
Data cutoff date: [YYYY-MM-DD]
Prepared by: [Name / Role]
Reviewed by: [Name / Role]
Approved by: [Name / Role]
This Statistical Analysis Plan (SAP) prespecifies the statistical methods for the study before final outcome analysis. The purpose is to reduce data-driven analytical decisions, improve reproducibility, and provide a transparent framework for estimating and interpreting epidemiologic associations or effects.
This template is designed primarily for observational epidemiologic studies, including:
Sections not relevant to a specific study should be marked Not Applicable (N/A) rather than silently omitted.
This SAP covers:
Provide a concise scientific and epidemiologic rationale.
Background:
[Describe the disease, exposure/intervention, public health relevance,
and prior evidence.]
Knowledge gap:
[Describe what remains uncertain and why the current study is
needed.]
Primary research question:
Among [target population], is
[exposure] associated with / causally related to
[outcome] over [follow-up period],
compared with [reference exposure]?
If the study is explicitly causal, define the causal contrast and the assumptions required for causal interpretation. If those assumptions are not supportable, results will be described as adjusted associations rather than causal effects.
To estimate the association/effect of [primary exposure] on [primary outcome] in [target population].
[Specify exploratory analyses. These analyses will be clearly distinguished from confirmatory analyses.]
For a two-sided primary comparison:
\[ H_0: \theta = \theta_0 \]
\[ H_A: \theta \neq \theta_0 \]
where \(\theta\) is the prespecified target estimand (for example, a hazard ratio, odds ratio, risk ratio, risk difference, mean difference, rate ratio, or marginal effect).
For ratio measures, \(\theta_0 = 1\). For difference measures, \(\theta_0 = 0\).
Design: [Prospective cohort / retrospective cohort / case-control / cross-sectional / nested case-control / case-cohort / registry-based / other]
Describe:
Specify each source and its role:
| Data source | Population coverage | Key variables | Date range | Linkage method |
|---|---|---|---|---|
| [Source 1] | [ ] | [ ] | [ ] | [ ] |
| [Source 2] | [ ] | [ ] | [ ] | [ ] |
Potential sources include cohort databases, EHRs, claims, disease registries, surveys, laboratory systems, pharmacy data, vital records, and linked administrative data.
The primary analysis will use data available through [date]. Any records received after this cutoff will not be incorporated into the primary analysis unless the SAP is formally amended before unblinded/final outcome analysis.
Define the population to which inference is intended to generalize.
Define the population represented by the data source.
Eligibility will be determined without reference to post-index outcomes whenever possible.
Provide a reproducible flow from the source population to the final analytic cohort.
Recommended flow:
Where appropriate, define:
Time zero is [definition].
Exposure classification and baseline covariates must be defined relative to this time point. Temporal alignment should minimize immortal-time bias and prevent use of post-outcome information as baseline information.
Primary exposure: [Definition]
Specify:
Example coding:
| Exposure category | Operational definition | Analysis code |
|---|---|---|
| Reference | [ ] | 0 |
| Exposed | [ ] | 1 |
For continuous exposures, specify the unit of interpretation (e.g., per 5 units, per SD, log-transformed) and whether nonlinearity will be modeled using restricted cubic splines or another prespecified method.
Primary outcome: [Definition]
Specify:
Secondary outcomes: [List]
Follow-up begins at [time zero] and ends at the earliest of:
Censoring assumptions and potential informative censoring will be evaluated as described in the sensitivity-analysis section.
For each key objective, define the target estimand explicitly.
Population: [Target population]
Exposure contrast: [Exposure A vs B / per unit
increase]
Outcome: [Outcome]
Time horizon: [e.g., 5 years]
Summary measure: [HR / RR / OR / RD / mean difference /
RMST difference / incidence rate ratio]
Handling of competing events: [strategy]
Handling of exposure changes/intercurrent events:
[strategy]
Covariate-adjustment target: [conditional or marginal
estimand]
The analysis should distinguish clearly between a conditional regression coefficient and a marginal population-level effect when those quantities differ.
CVD risk assuming everyone has diabetes CVD risk assuming no one has diabetes
[List each estimand using the same structure.]
Candidate confounders will be identified primarily using:
Variables will not be selected solely because their univariable p-values are statistically significant.
Create a covariate specification table:
| Covariate | Role | Definition/window | Coding | Missing-data approach | Primary model? |
|---|---|---|---|---|---|
| Age | Confounder | Baseline | Continuous | [ ] | Yes |
| Sex | Confounder | Baseline | Categorical | [ ] | Yes |
| [ ] | [ ] | [ ] | [ ] | [ ] | [ ] |
Unless required by the estimand, the primary adjustment set should generally avoid:
If a variable has uncertain causal status, alternative adjustment sets may be examined in sensitivity analyses.
Continuous covariates will be modeled as [linear / spline / categorized based on clinically justified cut points].
When restricted cubic splines are used, specify the number and location of knots before outcome analysis.
Before inferential modeling, perform structured data-quality checks.
For continuous variables, assess:
Values determined to be data errors will be corrected only using documented source information. Otherwise, values will be retained, set to missing, truncated, or excluded according to prespecified rules.
Summarize missingness by:
Assess known limitations in measurement sensitivity, specificity, coding changes, diagnostic algorithms, or self-report. Quantitative bias analysis may be considered when suitable external information is available.
Summarize follow-up duration, censoring reasons, loss to follow-up, and person-time by exposure group.
Report counts through each cohort-construction step.
Summarize baseline characteristics overall and by exposure group.
For continuous variables:
For categorical variables:
When comparing exposure groups, standardized mean differences (SMDs) will be presented when useful. SMDs are descriptive and will not determine confounder inclusion by themselves.
An absolute SMD of 0.10 may be used as a practical indicator of imbalance/balance, particularly for propensity-score diagnostics, but it is not a causal criterion for covariate selection.
Depending on design and outcome type, summarize:
Unless otherwise specified, statistical tests will be two-sided with \(\alpha = 0.05\). Estimates will be presented with 95% confidence intervals.
Primary interpretation will emphasize:
rather than statistical significance alone.
Specify presentation precision before final output production.
Recommended defaults:
<0.001 when
smaller.Analyses will be conducted using R version [x.x.x] and/or SAS version [x.x]. Key package/library versions will be archived for reproducibility.
The primary method will be selected according to the outcome scale and target estimand.
Possible primary models:
Generic model:
\[ g\{E(Y_i)\} = \beta_0 + \beta_1 X_i + \boldsymbol{\beta}_2^T \mathbf{C}_i \]
where \(X_i\) is the exposure and \(\mathbf{C}_i\) is the prespecified adjustment set.
Possible models:
Report adjusted mean differences or marginal contrasts as appropriate.
Possible models:
Include log person-time as an offset for incidence-rate analyses.
The default primary model may be Cox proportional hazards regression:
\[ h(t|X,C)=h_0(t)\exp\{\beta_1X+\boldsymbol{\beta}_2^T C\} \]
Report hazard ratios with 95% confidence intervals.
If the proportional hazards assumption is not appropriate for the exposure effect, prespecified alternatives may include:
Possible models:
Specify:
Use unconditional logistic regression unless matching requires conditional logistic regression.
Odds ratios will be interpreted in relation to the sampling design. Incidence-density sampling may permit the odds ratio to estimate an incidence rate ratio under appropriate assumptions.
Because temporality may be limited, results will generally be interpreted as prevalence associations rather than causal effects unless temporal ordering is otherwise established.
For the primary exposure-outcome relationship, report when appropriate:
Differences between crude and adjusted estimates may demonstrate the impact of adjustment but will not be used by themselves to identify true confounders.
Diagnostics will be appropriate to the selected model.
Evaluate:
cox.zph,Evaluate:
Evaluate:
Model diagnostics will inform prespecified alternative models or sensitivity analyses; they will not be used to conduct unrestricted data-driven model searching.
Propensity-score methods may be used as primary or sensitivity analyses when appropriate to the study objective.
Estimate:
\[ e(C)=P(X=1|C) \]
using prespecified baseline confounders.
Variables should be selected for confounding control, not solely based on statistical prediction of exposure.
Specify the estimand:
For ATE IPTW:
\[ w_i = \frac{X_i}{e(C_i)} + \frac{1-X_i}{1-e(C_i)} \]
Stabilized weights may be preferred for improved precision.
Assess:
Prespecify truncation/winsorization thresholds, for example the 1st/99th or 0.5th/99.5th percentiles, if trimming is planned.
Assess weighted SMDs and graphical diagnostics (e.g., Love plots). An absolute SMD < 0.10 may be considered acceptable measured balance.
Use weighted regression appropriate to the outcome with robust/sandwich standard errors as required.
Measured balance does not demonstrate absence of unmeasured confounding.
If exposure changes over time, specify:
When time-varying confounders are affected by prior exposure, conventional regression adjustment may introduce bias. Marginal structural models with inverse-probability weights may be used if prespecified and scientifically justified.
Report the amount and pattern of missing data for exposure, outcome, covariates, and follow-up variables.
Specify one primary strategy:
When MI is used:
Imputation should be consistent with the temporal structure of the study and should not inadvertently use future information in a way incompatible with the estimand.
When clinically relevant, evaluate departures from MAR using:
List all censoring mechanisms and justify non-informative censoring assumptions.
Potential approaches include inverse-probability-of-censoring weights (IPCW) or sensitivity analyses under plausible missing-outcome scenarios.
If competing events preclude the outcome, prespecify whether the scientific question targets:
Cause-specific Cox and Fine-Gray models answer different questions and should not be treated as interchangeable.
Prespecified effect modifiers may include:
Effect modification will preferably be evaluated by an interaction term rather than separate within-subgroup significance tests.
Example:
\[ g\{E(Y)\}=\beta_0+\beta_1X+\beta_2Z+\beta_3(X\times Z)+\cdots \]
Report stratum-specific estimates and the interaction p-value or confidence interval for the interaction contrast.
Subgroup analyses are generally considered supportive or exploratory unless explicitly designated as confirmatory.
For continuous or ordinal exposures, assess the prespecified form using one or more of:
If categories are used, cut points should preferably be prespecified or clinically motivated rather than chosen from observed outcome data.
If there is one primary exposure-outcome estimand, no multiplicity adjustment is needed for that single primary test.
When more than one confirmatory primary hypothesis is tested, specify a control strategy such as:
Unless otherwise prespecified, secondary and exploratory p-values will be interpreted descriptively with emphasis on effect sizes, uncertainty, and consistency rather than binary significance claims.
Sensitivity analyses should target specific assumptions rather than simply repeat the analysis using arbitrary alternatives.
Recommended categories include:
Compare, where appropriate:
Evaluate alternative truncation thresholds.
Examples:
Examples:
Use lag analyses or exclude events occurring shortly after baseline when scientifically appropriate and prespecified.
Compare complete-case, MI, and MNAR scenarios as appropriate.
Use IPCW or scenario analyses.
When appropriate, consider:
Evaluate alternative functional forms, distributions, correlation structures, or survival-model assumptions.
If available, prespecify negative-control outcomes or exposures that should not plausibly be caused by the exposure. Unexpected associations may indicate residual bias, measurement artifacts, or uncontrolled confounding.
Evaluate whether inclusion, retention, healthcare access, testing, referral, or data capture may depend jointly on exposure and outcome determinants.
If sampling weights are available, specify whether survey/design weights will be incorporated.
Compare the analytic cohort with the target/source population when possible and describe limits to transportability/generalizability.
For complex surveys, incorporate as appropriate:
Analyses must account for the design rather than treating observations as a simple random sample.
For clustering by site, family, physician, hospital, geographic area, or repeated person-level observations, use methods appropriate to the estimand, such as:
Specify the cluster level and correlation structure.
When sparse data may cause instability, prespecify alternatives such as:
Avoid fitting models with inadequate information relative to model complexity.
Outliers will not be removed solely because they affect statistical significance.
Potential influential observations will be:
The primary model and adjustment set will be prespecified.
Automated stepwise procedures should not be used to define the primary confounder set. If machine-learning methods are used for nuisance-model estimation, their role and tuning/cross-fitting procedures should be prespecified and separated from the causal estimand definition.
For high-dimensional EHR/claims settings, prespecified approaches may include:
These approaches do not replace careful specification of time zero, exposure, outcome, target population, causal structure, and estimand.
Observational studies may have a fixed available sample size. The SAP should state whether the study is:
Specify the anticipated number of exposed/unexposed participants, outcome events, and expected confidence-interval width for the primary effect.
Specify:
For weighting analyses, report effective sample size and recognize that extreme weights may substantially reduce precision.
Usually N/A for standard observational epidemiologic analyses.
If interim looks, sequential surveillance, or adaptive data collection are planned, specify stopping rules and error-control considerations before analysis begins.
Table 1. Cohort construction / participant
flow
Table 2. Baseline characteristics overall and by
exposure
Table 3. Missingness summary
Table 4. Outcome incidence / person-time by
exposure
Table 5. Crude and adjusted primary effect
estimates
Table 6. Secondary outcome analyses
Table 7. Prespecified subgroup / interaction
analyses
Table 8. Propensity-score balance diagnostics
Table 9. Sensitivity-analysis comparison
Figure 1. Study cohort flow diagram
Figure 2. Exposure distribution
Figure 3. Kaplan-Meier / cumulative incidence curve, if
applicable
Figure 4. Propensity-score overlap plot, if
applicable
Figure 5. Love plot / covariate balance, if
applicable
Figure 6. Forest plot of main and sensitivity
estimates
Figure 7. Restricted cubic spline exposure-response
curve, if applicable
The following code skeleton is intentionally generic and should be adapted to the specific study.
library(dplyr)
library(tidyr)
library(ggplot2)
library(tableone)
library(survival)
library(cobalt)
library(survey)
library(splines)
library(broom)
analysis_data <- raw_data %>%
filter(
eligibility_flag == 1,
!is.na(exposure),
index_date >= study_start,
index_date <= study_end
)
# Duplicate IDs
sum(duplicated(analysis_data$id))
# Missingness
colSums(is.na(analysis_data))
# Distributions
summary(analysis_data)
vars <- c(
"age",
"sex",
"bmi",
"comorbidity_1",
"comorbidity_2"
)
tab1 <- CreateTableOne(
vars = vars,
strata = "exposure",
data = analysis_data,
test = FALSE
)
print(tab1, smd = TRUE)
fit_binary <- glm(
outcome ~ exposure + age + sex + bmi + confounder_1,
family = binomial(),
data = analysis_data
)
summary(fit_binary)
fit_continuous <- lm(
outcome ~ exposure + age + sex + bmi + confounder_1,
data = analysis_data
)
summary(fit_continuous)
fit_cox <- coxph(
Surv(followup_time, event) ~
exposure + age + sex + bmi + confounder_1,
data = analysis_data
)
summary(fit_cox)
# PH assumption
ph_test <- cox.zph(fit_cox)
print(ph_test)
plot(ph_test)
ps_fit <- glm(
exposure ~ age + sex + bmi + confounder_1 + confounder_2,
family = binomial(),
data = analysis_data
)
analysis_data <- analysis_data %>%
mutate(
ps = predict(ps_fit, type = "response"),
p_exp = mean(exposure == 1),
sw = ifelse(
exposure == 1,
p_exp / ps,
(1 - p_exp) / (1 - ps)
)
)
# Prespecified truncation example
limits <- quantile(analysis_data$sw, c(0.01, 0.99), na.rm = TRUE)
analysis_data <- analysis_data %>%
mutate(
sw_trim = pmin(pmax(sw, limits[1]), limits[2])
)
bal.tab(
exposure ~ age + sex + bmi + confounder_1 + confounder_2,
data = analysis_data,
weights = analysis_data$sw_trim,
method = "weighting",
un = TRUE,
thresholds = c(m = 0.10)
)
fit_iptw <- coxph(
Surv(followup_time, event) ~ exposure,
data = analysis_data,
weights = sw_trim,
robust = TRUE
)
summary(fit_iptw)
fit_spline <- coxph(
Surv(followup_time, event) ~
ns(exposure_continuous, df = 4) +
age + sex + bmi + confounder_1,
data = analysis_data
)
Create one summary table containing all major estimates with the same target contrast.
Example structure:
| Analysis | Estimand/measure | Estimate | 95% CI | Key assumption changed |
|---|---|---|---|---|
| Crude | HR | [ ] | [ ] | No adjustment |
| Primary adjusted | HR | [ ] | [ ] | Primary confounder set |
| IPTW | Marginal HR / target measure | [ ] | [ ] | PS weighting |
| Alternative adjustment | HR | [ ] | [ ] | Covariate set |
| Missing-data sensitivity | HR | [ ] | [ ] | Missingness assumption |
Consistency across analyses will be evaluated based on effect direction, magnitude, precision, and the assumptions each analysis changes—not solely on whether individual p-values cross 0.05.
The final interpretation will distinguish among:
Results will not automatically be described as causal merely because regression, propensity scores, or weighting were used.
At minimum, discuss:
Any material deviation from this SAP after finalization will be documented with:
Post hoc analyses will be clearly labeled exploratory.
The study team will maintain:
| Role | Name | Signature | Date |
|---|---|---|---|
| Lead Statistician | [ ] | [ ] | [ ] |
| Epidemiologist / PI | [ ] | [ ] | [ ] |
| Data Manager / Analyst | [ ] | [ ] | [ ] |
| Other Reviewer | [ ] | [ ] | [ ] |
Complete this table before final analysis.
| Item | Prespecified decision |
|---|---|
| Primary research question | [ ] |
| Study design | [ ] |
| Target population | [ ] |
| Time zero | [ ] |
| Primary exposure | [ ] |
| Exposure reference | [ ] |
| Primary outcome | [ ] |
| Follow-up end | [ ] |
| Primary estimand | [ ] |
| Primary effect measure | [ ] |
| Primary confounder set | [ ] |
| Primary statistical model | [ ] |
| Continuous-variable functional forms | [ ] |
| Missing-data method | [ ] |
| Censoring strategy | [ ] |
| Competing-risk strategy | [ ] |
| PS method, if any | [ ] |
| Weight truncation, if any | [ ] |
| Prespecified effect modifiers | [ ] |
| Multiplicity strategy | [ ] |
| Sensitivity analyses | [ ] |
| Software/version | [ ] |
| Data cutoff | [ ] |
| Outcome type | Common primary model | Typical effect measure | Common alternatives |
|---|---|---|---|
| Binary | Logistic / modified Poisson | OR / RR / RD | GEE, log-binomial |
| Continuous | Linear regression | Mean difference | GLM, robust regression |
| Count/rate | Poisson / negative binomial | Rate ratio | GEE, recurrent-event models |
| Time-to-event | Cox PH | HR | RMST, AFT, flexible parametric |
| Repeated continuous | LMM / GEE | Mean contrast | Marginal models |
| Repeated binary | GEE / GLMM | OR / RR | Marginal standardization |
| Competing risk | Cause-specific Cox / Fine-Gray | Cause-specific HR / subdistribution HR | CIF risk contrast |
| Assumption | Primary approach | Suggested sensitivity analysis |
|---|---|---|
| Confounder specification | DAG/subject-matter set | Expanded/reduced set |
| Missing covariates | MI or complete case | Alternative missingness model |
| Exposure misclassification | Main algorithm | Restrictive/high-specificity definition |
| Outcome misclassification | Main algorithm | Validated/high-specificity definition |
| Positivity | Primary PS approach | Trimming/overlap weighting |
| Informative censoring | Standard censoring | IPCW |
| Reverse causation | Main lag | Longer lag/exclude early events |
| PH assumption | Cox | Time-varying effect/RMST |
| Unmeasured confounding | Standard adjustment | E-value/QBA/negative controls |
Adjusted association wording:
“After adjustment for prespecified baseline covariates, the exposure was
associated with [higher/lower] [outcome] ([effect measure] = [estimate],
95% CI [lower, upper]).”
Causal wording when assumptions are explicitly
supported:
“Under the stated assumptions of exchangeability, consistency,
positivity, correct model specification, and adequate measurement of
exposure, outcome, and confounders, the estimate may be interpreted as
the causal effect of [exposure contrast] on [outcome].”
Cautious observational wording:
“Because this is an observational study, residual and unmeasured
confounding, measurement error, selection bias, and uncertainty in
temporal ordering may remain; therefore, the estimate should not be
interpreted as definitive proof of causality.”