The Core Decision: Document Ownership Between Protocol and Analysis Plan
When a clinical development team prepares a confirmatory trial, an essential operational question arises: what endpoint specifications belong in the clinical study protocol, and what must be locked in the statistical analysis plan (SAP) before breaking the blind? ICH E9 section 5.1 treats that split as a confirmatory-status rule: only analyses envisaged in the protocol, including amendments, can be regarded as confirmatory. Leaving endpoint derivations vague in the protocol therefore cannot be cured later by SAP-only text; conversely, attempting to compress full algorithmic coding, visit-window algorithms, and estimation procedures into the protocol forces administrative protocol amendments for routine computational refinements.
The fundamental division is straightforward: the protocol locks the clinical question and the design architecture, while the statistical analysis plan locks the executable mathematical and algorithmic derivation of each endpoint before anyone can inspect treatment assignments. The protocol establishes the trial's target population, primary and secondary objectives, observation schedules, and the target estimands that express the therapeutic questions. The SAP translates those objectives into unambiguous, programmable instructions so that analysis-set membership, endpoint derivations, and estimators are not left to post-unblinding judgement.
This publication addresses document ownership and pre-unblinding lock boundaries. It builds upon our earlier analyses of ICH E9(R1) estimand specification, estimand-aligned sensitivity analysis for missing data, multiplicity management across endpoint hierarchies, and ICH E6(R3) implementation controls. Here, we focus specifically on the document-level governance: which principal features must reside in the protocol to retain confirmatory status, which executable specifications the SAP must finalize before the blind is broken, and when an analytical revision requires a formal protocol amendment rather than an internal SAP version update.
Protocol Principal Features Versus SAP Elaboration
The formal relationship between the protocol and the SAP was codified in ICH E9 (Statistical Principles for Clinical Trials, 1998) and reinforced in ICH E6(R3) (Good Clinical Practice, consolidated Step 4, June 2026) and ICH E8(R1) (General Considerations for Clinical Studies, 2021). ICH E9 section 5.1 explicitly distinguishes the two instruments:
Protocol Principal Features: The protocol must describe the principal statistical features of the eventual analysis, including the confirmatory analysis of the primary variable(s) and the way in which anticipated analysis problems will be handled. ICH E8(R1) section 5.6 separately keeps analyses of primary and secondary endpoints that address key objectives, interim analyses, planned design adaptations, analytical methods for planned estimation and hypothesis tests, and a sample-size justification in the protocol. A testing hierarchy among multiple endpoints is a prospective-specification lock under FDA's 2022 multiple-endpoints guidance, not an ICH E9 5.1 checklist item.
Statistical Analysis Plan Elaboration: The SAP is defined in the ICH E9 glossary as a separate, more technical elaboration of those principal features. It provides detailed, reproducible procedures for executing the statistical analysis of primary and secondary variables and other clinical trial data.
ICH E6(R3) Principle 8.3 requires that the protocol and trial execution plans—including the SAP—be clear, concise, and operationally feasible. Section 3.16.2(a) of ICH E6(R3) states that the sponsor must develop a SAP consistent with the protocol that details the approach to data analysis, unless that approach is already sufficiently described in the protocol itself. In modern registration trials, the complexity of computerized outcome assessments, repeated-measures models, and covariate adjustments makes a standalone SAP indispensable.
ICH E8(R1) and the FDA's April 2022 implementation clarify the statistical components that must remain anchored in the protocol. The protocol should define the estimands under ICH E9(R1), specify the analyses of primary and secondary endpoints that address key study objectives (including planned interim analyses and design adaptations), outline the analytical methods for point estimation and hypothesis tests, and supply a rigorous sample-size justification. The separate SAP then elaborates these foundations into a locked operational specification.
| Trial Design Element | Protocol Anchor (Principal Feature) | SAP Execution Lock (Elaboration) | Governing Standard |
|---|---|---|---|
| Primary & Secondary Endpoints | Clinical concept, measurement instrument, nominal assessment schedule, and hierarchy of importance. | Algorithmic derivation formula, visit-window definitions, unit standardisation, composite scoring rules, and re-test resolution. | ICH E9 §2.2.2; Gamble Items 26a–c |
| Target Estimand | The five ICH E9(R1) A.3.3 attributes: treatment, population, variable, handling of intercurrent events, and population-level summary. | Exact mapping of each estimand attribute to analysis datasets, variable names, and algorithmic handling. | ICH E9(R1) §A.6; Kang et al. (2022) |
| Main Estimator & Analysis Model | General class of statistical model (e.g., MMRM, ANCOVA, proportional hazards) and primary hypothesis test. | Fully specified linear/generalised model: fixed effects, covariates, interaction terms, covariance matrix structure, and software algorithms. | ICH E9 §5.1; Gamble Items 27a–b |
| Sensitivity Analyses | Commitment to evaluate robustness of primary conclusions to underlying distributional and missingness assumptions. | Pre-specified alternative models evaluating the same estimand (e.g., tipping-point parameter grids, pattern-mixture models). | ICH E9(R1) §A.5.2; Gamble Item 27e |
| Analysis Populations | Broad definitions of analysis sets (e.g., Full Analysis Set based on intention-to-treat, Safety Set). | Unambiguous programmatic criteria for subject inclusion/exclusion, protocol deviation flags, and handling of baseline dropouts. | ICH E9 §5.2; ICH E6(R3) §3.16.2(d) |
| Missing Data Handling | Guiding principles for missing-data prevention and general modeling strategy. | Reporting of missing-data assumptions and the statistical methods that will handle missing observations (Gamble's example is multiple imputation), including any same-estimand sensitivity analyses that test those assumptions. | ICH E8(R1) §5.6; Gamble Item 28 |
| Multiplicity & Error Control | Overall familywise Type I error rate and general testing hierarchy or gatekeeping architecture. | Step-by-step testing sequence, alpha allocation formulas, boundary calculation parameters, and correlation assumptions. | FDA Multiplicity Guidance (2022); Gamble Item 17 |
When the SAP Must Be Finalised, and What a Blind Review May Change
The timeline for locking a SAP is anchored to blinding and data access. Under ICH E9 section 5.1, the SAP must be reviewed, updated as necessary following a blind review of trial data, and formally finalised before breaking the blind. Formal records should be kept of when the SAP was finalised and when the blind was subsequently broken, so that finalisation can be shown to have preceded unblinding.
A critical operational question during trial conduct is what changes may be introduced during the blind review of aggregate data. ICH E9 section 5.1 establishes a strict boundary:
Protocol Amendment Required: If blind review of aggregate data reveals an unexpected distribution, higher-than-expected pooled variance, severe violation of normality, or unfeasible visit windows that requires changing the principal statistical features stated in the protocol—such as altering the primary endpoint, revising sample size, or shifting the testing hierarchy—that change must be executed via a formal protocol amendment. Under ICH E9 section 2.2.2, redefinition of the primary variable after unblinding is almost always unacceptable; even prior to unblinding, changing a primary endpoint in the SAP without amending the protocol creates an irreconcilable inconsistency.
SAP Version Update Sufficient: If the blind review merely clarifies computational execution—such as specifying a transformation algorithm to stabilise variance, defining visit-window boundaries for off-schedule assessments, or refining table shells and SAS/R macro parameters—these adjustments belong in an updated, version-controlled SAP.
Trial masking dictates distinct deadlines under ICH E8(R1) and the FDA's April 2022 implementation:
Double-Blind Studies: The SAP must be finalised before treatment assignments are revealed to anyone involved in trial execution or analysis. Crucially, the planned analysis cannot be modified following an unblinded interim analysis, such as an interim review conducted for efficacy or futility by an independent Data Monitoring Committee (DMC).
Open-Label and Single-Blind Studies: Because treatment assignments are visible from enrollment, bias can infiltrate analytical choices during data collection. Therefore, the details of the primary and important secondary analyses should ideally be finalised before the first participant is randomised or allocated to treatment.
What happens if changes are contemplated after unblinding? ICH E6(R3) section 3.16.2(e) restricts deviations from the planned statistical analysis, and changes made to the data after unblinding, to exceptional, documented, and justified circumstances. Investigator authorisation and an audit trail apply to those data changes; both the data changes and the analysis deviations must appear in the clinical trial report. Protocol item B.10.5 requires the protocol to state that SAP deviations will be described and justified in that report. ICH E9 section 5.1 already withholds confirmatory status from analyses that were not envisaged in the protocol, including amendments.
flowchart TD
A["Protocol Finalisation\n(Principal Features and Target Estimands)"] --> B["Trial Execution and Data Collection\n(Treatment Masking Maintained)"]
B --> C["Blind Data Review\n(Aggregate Data, Blinded to Assignment)"]
C --> D{"Does Review Change\nPrincipal Features?"}
D -- "Yes: primary variable, hierarchy, or sample size" --> E["Formal Protocol Amendment"]
E --> F1["Update SAP to Match Amendment"]
D -- "No: visit windows, derivations, or coding" --> F2["Update SAP Version"]
F1 --> G["Date-Stamped SAP Finalisation\n(Recorded Before Blind Break)"]
F2 --> G
G --> H["Database Lock and Unblinding"]
H --> I["Execute Pre-Specified Analysis"]
H -. "Unplanned Deviation" .-> J["Documented Exceptional Deviation\n(Justified in CSR; Not Confirmatory)"]Per Endpoint, Lock the Aligned Estimator and Same-Estimand Sensitivity Analyses
The release of ICH E9(R1) (Addendum on Estimands and Sensitivity Analysis, Step 4, November 2019) resolved decades of conflation between clinical trial questions and mathematical estimators. ICH E9(R1) section A.6 articulates the precise document division:
The protocol should explicitly define the primary estimand corresponding to the primary trial objective, as well as any key secondary estimands intended to support regulatory claims. The protocol and the statistical analysis plan should then pre-specify the main estimator aligned with that estimand, together with a suitable sensitivity analysis to explore robustness to departures from the estimator's core assumptions.
Same-Estimand Sensitivity Analyses Versus Supplementary Analyses
A critical regulatory trap is confusing sensitivity analyses with supplementary analyses. ICH E9(R1) sections A.5.2 and A.5.3 enforce a clear distinction:
Same-Estimand Sensitivity Analysis (§A.5.2): A sensitivity analysis investigates the statistical assumptions of the main estimator while targeting the exact same estimand. Its sole objective is to verify whether the trial's conclusions are robust to deviations from those assumptions. For example, if a primary estimator for a continuous endpoint relies on a missing-at-random (MAR) assumption within a mixed-effects model for repeated measures (MMRM), a pre-specified tipping-point analysis or pattern-mixture model targeting the same treatment effect under missing-not-at-random (MNAR) departures constitutes a valid sensitivity analysis. The SAP must pre-specify the exact grids, delta shifts, and algorithms for these robustness tests.
Supplementary Analysis (§A.5.3): A supplementary analysis evaluates a different estimand or investigates the trial data under an alternative clinical perspective to build broader understanding (for instance, examining on-treatment adherence versus treatment policy). While valuable for exploratory insights, supplementary analyses play a subordinate role in regulatory decision-making and cannot substitute for a same-estimand sensitivity analysis.
In peer-reviewed methodological literature, Kang et al. (Clinical Trials, 2022) operationalised this framework by publishing an estimand-to-analysis table template. In this structure, the ESTIMAND column restates each attribute of the estimand (treatment, target population, variable, intercurrent events, and population-level summary), while the ANALYSIS column locks how each attribute will be handled with the collected data, including missing-outcome methods. Kang illustrated the template with a hypothetical paediatric HIV safety trial; it is not ICH text. Lynggaard et al. (Trials, 2022) similarly keep planned analyses for main estimands in the clinical study protocol and allow estimation of less important estimands to be deferred to the SAP. That option is a protocol-template recommendation, not an ICH permission to omit the primary estimand from the protocol.
What Gamble Requires the SAP to Specify for Each Outcome
While ICH E9 establishes regulatory principles, it does not provide an exhaustive per-endpoint document checklist. The recognized international benchmark for SAP content is the consensus guideline published by Gamble et al. (JAMA, 2017), developed through an extensive Delphi process and endorsed by the EQUATOR Network. Comprising 55 addressable elements under 32 numbered headings, Gamble et al. defines what an audit-ready SAP must detail beyond the protocol summary.
Gamble assumes the SAP is read in conjunction with a SPIRIT-consistent protocol and is applied to clean, validated analysis datasets. The guideline emphasizes that early authoring before data collection is optimal, and that the blind review represents the final permissible window for SAP amendments.
For each primary and secondary outcome assessment, the SAP must provide granular, reproducible specifications across several key Gamble items:
Outcome Definition, Timing, and Derivation (Items 26a–c): Items 26a–c require the SAP to list each primary and secondary outcome with timings and, if applicable, the order in which primary or key secondary endpoints will be tested; the specific measurement and units; and any calculation or transformation used to derive the outcome (Gamble's examples include change from baseline, a quality-of-life score, time to event, and a logarithm). Visit windows themselves are Gamble item 15: the time points at which outcomes are measured, including visit windows. Neither item invents a numeric window width that applies to every endpoint.
Statistical Analysis Methods and Presentation (Items 27a–b): Item 27a requires the analysis method and how treatment effects will be presented. Item 27b requires any covariate adjustment. ICH E9 section 5.7 adds that if one or more factors are used to stratify the design, it is appropriate to account for those factors in the analysis; that is an alignment choice to lock, not a universal covariate list.
Assumption Checks and Fallback Procedures (Items 27c–d): Item 27c requires the methods used for assumptions to be checked for the statistical methods. Item 27d requires alternative methods if distributional assumptions do not hold (Gamble's examples are normality and proportional hazards). Named diagnostics such as residual plots or a score test for proportional hazards are illustrations of what 27c can lock, not a required battery. Crucially, the plan must pre-define alternative fallback methods to be executed if distributional assumptions fail (for instance, specifying a pre-planned Wilcoxon rank-sum test or robust Huber-White sandwich estimator if normality is violated). Selecting a non-parametric test only after observing a non-significant parametric result is an unacceptable post-hoc practice.
Pre-Specified Sensitivity and Subgroup Analyses (Items 27e–f): Item 27e requires planned sensitivity analyses for each outcome where applicable. Item 27f requires planned subgroup analyses, including how subgroups are defined. Any numeric cut, such as an age split, is a trial-specific definition to lock—not a universal threshold.
Missing Data Handling (Item 28): Item 28 requires reporting of assumptions and the statistical methods used to handle missing data (Gamble's example is multiple imputation). The SAP should name the missingness assumptions of the main estimator and the same-estimand sensitivity analyses that test them; it does not prefer a single imputation algorithm.
Multiplicity and Analysis Populations (Items 17 and 20): The SAP must document the rationale for any multiplicity adjustment and how Type I error is to be controlled, alongside definitions of each analysis population (Gamble's examples include intention to treat, per protocol, complete case, and safety).
| Gamble 2017 Item | Required SAP Specification | Execution Hazard If Omitted | Auditable Deliverable |
|---|---|---|---|
| Item 15 (timing of outcome assessments) | Time points at which outcomes are measured, including visit windows, and rules for unscheduled or duplicate records. | Visit mapping or re-test selection decided after unblinding. | Protocol-referenced visit windows, locked as executable mapping rules in the SAP. |
| Item 26a (outcomes and timings) | Each primary and secondary outcome with timings and, if applicable, the order in which they will be tested. | Testing order or outcome timing chosen after seeing results. | Per-outcome definition with testing order if a hierarchy applies. |
| Item 26b (measurement and units) | Standardised measurement units, international conversion factors, and boundary validation limits. | Mismatched laboratory units corrupting pooled effect estimation. | Unit standardisation specifications in data derivation rules. |
| Item 26c (derivations) | Explicit algebraic derivation formulas, subscale scoring algorithms, and missing-item imputation thresholds. | Discrepant score derivations between programming teams and unblinded audits. | Validated software macro specifications and derivation formulas. |
| Items 27a–b (method and covariates) | Exact model formulation, treatment contrast coding, baseline covariates, and stratification factor adjustments. | Post-hoc covariate dredging or selective adjustment to achieve statistical significance. | Full mathematical model formula with defined fixed and random effects. |
| Items 27c–d (assumption checks and fallbacks) | Formal assumption test criteria and pre-specified alternative non-parametric or robust models. | Opportunistic switching between parametric and non-parametric tests post-unblinding. | Binary decision rules specifying when fallback models execute. |
| Item 27e (sensitivity analyses) | Pre-specified same-estimand robustness models testing departures from estimator assumptions. | Same-estimand robustness cannot be demonstrated from an analysis that was not pre-specified. | Algorithmic parameters for tipping-point or stress-testing analyses. |
| Item 28 (missing data) | Reporting of missing-data assumptions and the statistical methods that will handle missing observations. | Arbitrary post-hoc imputation choices that artificially reduce standard errors. | Named missing-data assumptions and methods, plus same-estimand sensitivity analyses that test them. |
| Item 17 (multiplicity) | Description and rationale for any multiplicity adjustment and how Type I error is to be controlled. | Additional analyses after unmasking reintroduce an uncontrolled multiplicity problem. | Pre-specified error-control procedure aligned to the protocol hierarchy. |
| Item 20 (analysis populations) | Definition of analysis populations (Gamble examples: intention to treat, per protocol, complete case, safety). | Analysis-set membership decided from unblinded listings rather than pre-defined rules. | Programmable inclusion and exclusion criteria, with ICH E9(R1) caution that a per-protocol set may not align to a relevant estimand. |
Analysis-Set Membership and Traceable Derivations
Decisions regarding which participants and observations are included in an analysis set cannot be left to post-unblinding discretion. Under ICH E9 section 5.2, decisions concerning the analysis set should be guided by minimising bias and avoiding inflation of Type I error. Section 5.2.1 defines the full analysis set as the set as complete as possible and as close as possible to the intention-to-treat ideal of including all randomised subjects. Failure to satisfy major entry criteria, failure to take at least one dose of trial medication, and lack of any data post randomisation are listed as limited circumstances that might lead to excluding a randomised subject; such exclusions should always be justified. Those circumstances are not a default membership rule, and lack of post-randomisation data is not the same as lack of a baseline record.
ICH E6(R3) section 3.16.2(d) requires that criteria for inclusion or exclusion of trial participants from any analysis set be pre-defined in the protocol or the SAP, and that the rationale for excluding any participant or data point be clearly described and documented. ICH E9 section 5.2.2 separately requires per-protocol criteria to be defined and documented before breaking the blind. Protocol item B.10.4 is a protocol requirement: the selection of participants in the planned analyses, a description of the statistical methods, and procedures for handling intercurrent events and accounting for missing, unused, and spurious data, aligned with the target estimands when defined.
A critical modern pitfall involves the conventional Per-Protocol Set (PPS). In historical practice, biostatisticians routinely copied a default 'FAS plus PPS' analysis structure into the SAP. However, ICH E9(R1) explicitly warns that it is frequently impossible to construct a well-defined causal estimand to which a standard per-protocol analysis corresponds. Excluding subjects who experienced protocol deviations or treatment discontinuation often introduces severe confounding and selection bias. Consequently, modern SAPs must avoid unreflective per-protocol exclusions, relying instead on estimand-aligned intercurrent event strategies (such as hypothetical or principal stratum strategies) pre-specified in the protocol.
Additionally, ICH E6(R3) section 3.16.2(c) enforces end-to-end data integrity: all data derivations and transformations must be traceable from raw electronically captured instruments (eCOA, EDC, sensors) through to analysis-ready datasets (e.g., CDISC ADaM standards). The SAP serves as the design blue-print against which statistical programming verification and independent double-programming are audited.
Multiplicity Controls and Adaptive Design Locks
A central purpose of the SAP is to preserve familywise error rate (FWER) control across multiple endpoints. Under the FDA's October 2022 final guidance, Multiple Endpoints in Clinical Trials, prospective specification is mandatory. The guidance emphasizes that to control multiplicity, sponsors must prospectively specify all planned endpoints, assessment time points, analysis populations, treatment doses, and analytical procedures. The statistical analysis plan should not be changed after unmasking treatment assignments and performing statistical analyses.
The FDA guidance cautions that introducing post-unblinding analyses or re-ordering testing sequences reintroduces uncontrolled multiplicity problems. While the sponsor has discretion to choose among various multiplicity adjustment strategies—such as fixed-sequence testing, fallback procedures, Bonferroni-based gatekeeping, or graphical error-splitting methods—the procedure, alpha allocation, and passing criteria must be prospectively specified in the SAP before unmasking. This article does not present one procedure as universally preferred.
In trials incorporating adaptive designs, documentation requirements are even more stringent. The FDA's November 2019 guidance, Adaptive Designs for Clinical Trials of Drugs and Biologics, requires that documentation submitted to the Agency prior to trial initiation describe the adaptation plan in addition to typical ICH E9 protocol and SAP components. That written documentation may be included in the protocol and/or in a SAP, a DMC charter, or an adaptation committee charter, but all important information should be submitted during the design stage so the review division has time to comment before the trial starts. Adequately pre-specified adaptations based on non-comparative (blinded) data generally have no effect or a limited effect on Type I error; comparative interim analyses require additional integrity controls. Unplanned endpoint changes after seeing unblinded results are outside this SAP-lock article.
What This Lock Does Not Claim: Methodological Limits and Non-Claims
To maintain scientific and regulatory rigor, clinical development teams must recognize what a locked statistical analysis plan accomplishes—and what it does not:
Regulatory Guidance Is Not Statutory Approval: ICH guidelines (E9, E9(R1), E8(R1), E6(R3)) and FDA guidance documents represent current agency thinking and non-binding recommendations. Locking a SAP according to these principles demonstrates methodological pre-specification; it does not constitute FDA or EMA approval of the endpoint, acceptance of the surrogate validity, or clearance of labeling claims.
No Single Estimator Is Globally Endorsed: ICH E9(R1) intentionally does not designate MMRM, multiple imputation, proportional hazards, or tipping-point analysis as universally preferred techniques. Every estimator must be scientifically justified based on the specific disease indication, instrument measurement properties, and estimand attributes.
SAP Pre-Specification Cannot Salvage Protocol Deficiencies: Under ICH E9 section 5.1, only analyses envisaged in the protocol (including formal protocol amendments) can support confirmatory conclusions. Defining a brand-new primary endpoint solely in the SAP without an approved protocol amendment will result in regulatory rejection of confirmatory claims.
Academic Checklists Are Guidelines, Not Regulations: Gamble et al. (JAMA 2017) and related templates (Kang et al. 2022, Lynggaard et al. 2022) are consensus reporting standards developed primarily for later-phase trials. While they represent industry best practice, they are not statutory requirements. For early-phase trials, biostatisticians should consider extensions such as Homer et al. (BMJ 2022), which adapts the Gamble framework for Phase I and non-randomised Phase II studies.
Pre-Specification Prohibits Post-Hoc Adaptation: Prospectively planned adaptations under FDA's 2019 adaptive design guidance require rigorous integrity firewalls and pre-specified algorithms. Pre-specification does not authorize ad-hoc modifications to endpoint definitions following unblinded interim inspections.
By maintaining a rigorous, auditable separation between protocol-level principal features and SAP-level execution locks, clinical trial sponsors protect the statistical validity of their endpoints and ensure that study results remain fully interpretable and inspection-ready under international regulatory scrutiny.
