Estimands & Statistics

Multiple Endpoints in Clinical Trials: Hierarchy, Co-Primaries, and Error Control

A rigorous analysis of multiplicity management in clinical trials, comparing FDA's 2022 final guidance and EMA's 2002 Points to consider across co-primaries, hierarchies, and gatekeeping.

· · 9 min read

Editorial still life of three stacked unmarked study folios tied with an oxblood cord on parchment beside muted navy cloth

Scenario Resolution: To preserve statistical rigor and regulatory interpretability across multiple clinical endpoints, sponsors must establish an explicit error-control hierarchy prior to study unblinding. Endpoints must be partitioned into primary, secondary, and exploratory families. Both FDA's October 2022 guidance and EMA's 2002 Points to Consider (CPMP/EWP/908/99) require controlling family-wise Type I error in the strong sense across primary and secondary claims. While co-primary endpoints require no alpha splitting, sponsors must increase sample size to compensate for substantial power loss. Secondary endpoints cannot yield confirmatory labelling claims based on nominal p-values unless pre-specified within a valid gatekeeping or fallback sequence.

Three texts, three regulatory statuses

Statistical multiplicity in clinical trials is governed by distinct regulatory frameworks across the United States and the European Union. Sponsors designing multinational confirmatory trials must navigate three distinct regulatory texts and their precise legal statuses:

  1. 1. FDA Final Guidance (October 2022): Titled Multiple Endpoints in Clinical Trials (docket FDA-2016-D-4460; content current as of 16 April 2024). This document represents FDA's finalized thinking on endpoint families, co-primaries, secondary claim hierarchies, and statistical testing methodologies, formally replacing the January 2017 draft. It establishes practical guidelines for managing Type I error across diverse clinical development settings, from simple two-arm studies to complex multi-dose, multi-endpoint factorial designs.

  2. 2. EMA Points to Consider (September 2002): Titled Points to consider on multiplicity issues in clinical trials (CPMP/EWP/908/99), adopted by the Committee for Proprietary Medicinal Products (CPMP) in September 2002, with an overseas effective date of 1 April 2003 recorded by adopting regulators such as TGA. EMA still lists this Points to consider as the current effective version. It emphasizes strong family-wise Type I error control, including closed testing procedures, and hierarchical testing. It is scientific guidance, not a statute.

  3. 3. EMA Draft Guideline (2017): Titled Draft guideline on multiplicity issues in clinical trials (EMA/CHMP/44762/2017). While intended to replace the 2002 Points to Consider, consultation closed in June 2017 and the document remains classified as a draft under EMA's document history. The draft explored advanced graphical testing methods and adaptive group sequential designs, but because it was never formally adopted as a Step 5 scientific guideline, sponsors must not cite this draft as operative EU guidance.

Primary, secondary, and exploratory families

FDA's October 2022 guidance establishes a rigorous tripartite classification for clinical trial endpoints, establishing how alpha must be partitioned across the study architecture:

  • Primary Endpoint Family: The outcome measure or measures established to demonstrate statistically persuasive evidence of clinical efficacy to support drug approval. In major disease indications, this may consist of a single clinical endpoint (such as overall survival in oncology), co-primary endpoints (such as cognitive and functional co-primaries in Alzheimer's disease), or alternative multiple primary endpoints. Failure to achieve statistical significance on the primary family terminates confirmatory hypothesis testing across the trial.

  • Secondary Endpoint Family: Selected outcome measures intended to support additional promotional or labelling claims regarding clinical benefits, symptom improvement, disease modification, or secondary indications. Secondary endpoints can only be formally evaluated for confirmatory claims if the primary family succeeds. Typical examples include key organ-specific biomarkers, health-related quality of life measures, or secondary clinical symptom scores.

  • Exploratory Endpoint Family: Hypothesis-generating variables, novel digital health sensor streams, and exploratory patient-reported scales. Exploratory endpoints do not require statistical multiplicity adjustment because they are not intended to support regulatory approval or promotional claims. However, sponsors cannot repurpose an exploratory endpoint into a confirmatory claim post-unblinding.

graph TD
    A["Overall Type I Error Budget: alpha = 0.05"] --> B["Primary Endpoint Family: Primary Test"]
    B -->|Primary Test Passes p < alpha| C["Secondary Endpoint Family: Hierarchical Alpha Propagation"]
    B -->|Primary Test Fails| D["Confirmatory Testing Halts: Secondaries Exploratory Only"]
    C --> E["Confirmatory Labelling Claim Licensure"]
    F["Exploratory Endpoint Family"] --> G["Internal Learning & Hypothesis Generation: No Claim Licensure"]
The Family-Wise Error Rate (FWER) hierarchy across primary, secondary, and exploratory endpoint families.

The central mandate of multiplicity control is that the overall family-wise error rate (FWER) must be controlled in the strong sense across the primary and secondary families collectively. Controlling Type I error in the strong sense means that the probability of making at least one false positive claim is bounded by alpha (typically 0.05 two-sided) under any configuration of true and false null hypotheses, rather than merely under the global intersection null hypothesis (weak control).

Co-primaries versus multiple primaries

A frequent source of trial failure is confusing co-primary endpoints with multiple primary endpoints. These two designs have diametrically opposed consequences for Type I error, statistical power, and sample size calculations:

Design DimensionCo-Primary EndpointsMultiple Primary Endpoints
Core RequirementAll specified endpoints must independently achieve p < 0.05Any single endpoint achieving significance constitutes trial success
Type I Error ImpactNo inflation; intersection-union principle is conservativeInflates Type I error unless alpha is split or otherwise controlled
Alpha AdjustmentNone required (tested at full nominal alpha)Mandatory adjustment (Bonferroni, Holm, Hochberg, etc.)
Type II Error ImpactSubstantial power loss; joint power is product of marginal powersPower increases for achieving at least one success
Typical Clinical ApplicationAlzheimer's disease (Cognition + Global Function); NASH (Fibrosis + Resolution)Migraine relief (Pain freedom OR Most bothersome symptom); Oncology (PFS OR OS)

The Co-Primary Joint Power Penalty and Correlation Impact

When two independent endpoints are designated as co-primaries, each powered individually at 80% (beta = 0.20), the joint probability of succeeding on both endpoints drops precipitously:
P(Success on Both) = P(Endpoint 1) x P(Endpoint 2) = 0.80 x 0.80 = 0.64 (64%)

As FDA's 2022 guidance highlights, to preserve an overall trial power of approximately 80% across two independent co-primary endpoints, each individual endpoint must be powered at approximately 90% (since 0.90 x 0.90 = 0.81). FDA therefore recommends a larger sample size so that the individual endpoints can be powered at approximately 90% when independence is a reasonable working assumption. The calculation differs if the endpoints are highly positively correlated or if power is not equal for each endpoint. Dual histological co-primaries used in metabolic-associated steatohepatitis development programs illustrate why that Type II penalty must be budgeted before FPI; this article does not estimate an industry failure rate for those trials.

When endpoints exhibit positive correlation, FDA notes that the joint-power calculation would be different. The October 2022 guidance does not publish rho-specific joint-power tables. Do not assume a correlation large enough to rescue an 80%/80% co-primary design without protocol-specific empirical support; if the true correlation is lower than assumed, the trial is under-powered for the all-must-succeed claim.

When a secondary endpoint can support a claim

Regulatory authorities will not permit a secondary endpoint finding to enter product labelling or promotional copy unless it is protected by a pre-specified, methodologically sound multiplicity adjustment strategy. FDA's guidance appendix and EMA's Points to Consider detail several acceptable frameworks:

  1. Fixed Sequence (Hierarchical) Testing: Hypotheses are ordered in a predefined clinical hierarchy (H1 -> H2 -> H3). H1 is tested at alpha = 0.05. If H1 is rejected, H2 is tested at alpha = 0.05. The first non-rejected hypothesis terminates formal testing, and all subsequent endpoints are relegated to exploratory status. This is the simplest and most widely accepted gatekeeping method.

  2. Step-Down Methods (Holm & Hochberg): Holm's step-down procedure orders p-values from smallest to largest and tests against adjusted thresholds alpha/(k - i + 1). It controls FWER strongly under any dependency structure. Hochberg's step-up procedure is more powerful but requires non-negative dependence (PRDS).

  3. The Fallback Procedure (Wiens 2003): Allows assigning fixed initial alpha weights to ordered hypotheses (e.g., alpha_1 = 0.03, alpha_2 = 0.02). If H1 is rejected, its alpha is passed to H2; if H1 fails, H2 can still be tested at its initial alpha_2. This prevents a complete testing halt if an early secondary endpoint narrowly misses significance.

  4. Serial Gatekeeping: All primary endpoints must pass before any secondary endpoint can be evaluated. If all primary tests succeed, the full alpha is transferred to the secondary family.

  5. Parallel Gatekeeping: Permits alpha to pass to the secondary family if at least one primary endpoint succeeds, using structured alpha-allocation algorithms (e.g., Dmitrienko-Tamhane-Wiens procedures).

  6. Graphical Approaches (Maurer & Bretz): A flexible framework visualizing hypotheses as nodes and alpha transfer weights as directed edges, ensuring strict strong-sense FWER control while allowing multidirectional alpha recycling upon hypothesis rejection.

Composites, components, and method choice

Composite endpoints—such as Major Adverse Cardiovascular Events (MACE: cardiovascular death, nonfatal myocardial infarction, nonfatal stroke)—are widely utilized in cardiovascular, oncology, and autoimmune trials to evaluate overall clinical impact while avoiding multiple testing penalties on the primary efficacy test.

However, sponsors frequently fall into the composite component trap. Demonstrating a statistically significant reduction in the composite endpoint does not automatically license promotional or labelling claims for individual component events. If a sponsor seeks a specific claim that a drug reduces cardiovascular mortality, that component must be formally pre-specified within the secondary error-control hierarchy.

Furthermore, FDA and EMA evaluators scrutinize composite endpoints to ensure that the aggregate treatment effect is not driven entirely by less severe components (such as elective revascularization or hospitalization) while critical components (such as mortality) show neutral or unfavorable trends. If a component is itself intended as a demonstrated effect, it should be pre-specified with multiplicity control. Ranking methods such as a win ratio can be used to order clinically dominant events, but neither FDA's 2022 guidance nor EMA's 2002 Points to consider designates a win ratio as a default or increasingly required composite analysis.

By embedding rigorous multiplicity control into the initial statistical design, clinical development sponsors preserve the evidentiary integrity of secondary endpoints, avoid false-positive regulatory rejections, and secure robust, label-supporting confirmatory claims.