The Core Decision: Score Invariance Across Modes vs. Device Logistics
When clinical operations and development teams evaluate electronic clinical outcome assessments (eCOA), executive discussions frequently gravitate toward procurement and hardware logistics: whether the sponsor should provision dedicated handheld smartphones, deploy an in-clinic tablet fleet, or implement a Bring Your Own Device (BYOD) model where participants download a study application onto their personal phones. While hardware deployment introduces serious budgetary, regulatory, and supply-chain considerations—as examined in our eCOA device strategy evaluation of BYOD, provisioned, and hybrid models—device sourcing is fundamentally distinct from the psychometric question of measurement comparability.
The central methodological question confronting biostatisticians and clinical trialists is whether a patient-reported outcome measure (PROM) preserves its exact clinical meaning, construct validity, and scoring distribution when completed across different collection modalities. If half of a trial cohort records symptom severity on personal smartphones while the remainder uses clinic tablets or web portals—or if a patient transitions between modalities midway through a 52-week confirmatory trial—can those data be pooled into a single primary analysis without introducing systematic response bias or inflating error variance?
Crucially, mode comparability must also be separated from neighboring regulatory and technical requirements. Establishing measurement comparability across modes is distinct from verifying sensor-derived digital measures in digital health technologies, qualifying computerized systems under risk-based computerized-system validation for endpoint capture, or demonstrating electronic record integrity under 21 CFR Part 11 audit trail controls. A system can be technically secure, fully encrypted, and Part 11 compliant, yet still fail to yield valid trial endpoints if an awkward screen migration distorts patient cognition or response behavior.
Evolution of Standards: From Mandatory Equivalence to Conditional Clearance
For more than fifteen years, sponsors often treated every paper-to-electronic migration as if it required a new equivalence study. Tracing the ISPOR reports shows why the current public decision rule is conditional rather than automatic testing or automatic mixing.
Coons et al. (2009): The Original Modification Ladder
The 2009 ISPOR ePRO Good Research Practices Task Force report (Coons et al.) established the foundational framework governing paper-to-electronic adaptations. Grounded in precautionary principles, the task force defined measurement equivalence as the comparability of psychometric properties across modes, categorizing adaptations across three distinct levels of modification:
Minor Modifications: No change in content or meaning: nonsubstantive instruction changes (for example, circling a response versus touching it on a screen) and minor format changes such as one item per screen rather than multiple items on a page. Coons et al. concluded that cognitive debriefing and usability testing may be sufficient because a substantial body of evidence already suggested that psychometric properties would hold.
Moderate Modifications: Changes that, on the then-available literature, could not be justified as minor and might alter content or meaning: more significant presentation changes, or a change in administration involving different cognitive processes (the usual example is paper [visual] to interactive voice response [aural]). Coons et al. called for quantitative equivalence testing plus usability testing, not a change of response scale type such as converting a visual analogue scale into a numeric rating scale.
Substantial Modifications: Changes with no existing empirical support that clearly alter content or meaning, including substantial changes in item wording or response options. Those cases were treated as requiring full psychometric testing in addition to usability testing. Changing recall period is still outside a mere mode migration; FDA's 2025 PFDD Guidance 3 treats a 1-day to 7-day recall change as potentially creating a new measure.
Eremenco et al. (2014): Strict Restrictions on Mixed Modes
In 2014, the ISPOR PRO Mixed Modes Task Force (Eremenco et al.) addressed clinical trials that combined multiple data collection methods within a single protocol. Written primarily for registrational studies intended to support promotional labeling claims, the 2014 report took a restrictive posture. In the absence of documented measurement equivalence, the task force strongly recommended conducting a prospective quantitative equivalence study before permitting mode mixing.
Furthermore, Eremenco et al. strongly discouraged mixing paper and electronic field-based instruments. If mixing proceeds, the 2014 report said data collection should be planned at the country level or higher, ad hoc mixing by sites or individual subjects should be minimized, and the statistical analysis plan must address mixing and evaluate whether data can be pooled before evaluating treatment effects.
O'Donohoe et al. (2023): The Two-Condition Decision Rule
By 2023, a large accumulated literature showed generally good paper-versus-electronic agreement for many PROMs, with important caveats about heterogeneity and about which modifications were actually studied. An updated ISPOR Task Force published new recommendations (O'Donohoe et al., Value in Health, May 2023) that fundamentally revised the 2009 and 2014 positions.
Rather than mandating routine testing for every adaptation, the 2023 Task Force established a unified, evidence-based decision principle:
"In cases where sufficient evidence of measurement comparability exists and best practices for faithful migration are followed, this task force concludes that further testing of measurement comparability among the data collection modes is unnecessary, including cases of mixing modes within clinical trials such as bring your own device designs."
McLeod and Rockwood (2023), in an accompanying editorial, described the update as opening the gates for faithful migration. The operational test is cumulative: existing evidence must satisfactorily show that the mode change did not affect measurement properties, and best practices for a faithful migration must have been followed. If both conditions are met, the Task Force concludes that further mode-comparability testing is unnecessary, including some mixed-mode and BYOD designs. If either condition fails, additional testing—from cognitive interviewing and usability testing through quantitative comparability testing—should be considered, depending on the measure and its intended use. ISPOR Good Practices reports are consensus methods, not FDA or EMA regulations.
| Framework & Year | Core Philosophy | Empirical Testing Triggers | Stance on Mixed Modes & BYOD |
|---|---|---|---|
| ISPOR ePRO (Coons 2009) | Precautionary three-tier modification ladder | Mandatory CD/UT for minor; quantitative crossover trial for moderate modifications | Paper-to-electronic migration focus; mixed-mode pooling was not the report's primary job |
| ISPOR Mixed Modes (Eremenco 2014) | Rigorous proof of poolability required | Quantitative equivalence study strongly recommended prior to pooling scores | Strongly discouraged paper + electronic field diaries; required SAP poolability rules |
| ISPOR Task Force (O'Donohoe 2023) | Evidence-based conditional clearance | Testing unnecessary if existing evidence exists AND faithful migration is documented | Permits electronic mixing and BYOD conditionally; paper remains highest-risk mode |
| C-Path eCOA Consortium (Mowlem 2024) | Operationalized Class 1 (comparability) vs Class 2 (usability) design standards | Not considered necessary for common paper-to-electronic scale types if at least Class 1 practices are followed; mixed-mode and BYOD still also need the 2023 evidence condition | When Class 1 practices and the 2023 evidence condition are both met, further comparability testing is not considered necessary, including mixed-mode and BYOD designs |
| FDA PFDD Guidance 3 (Final Oct 2025) | Current thinking on small-to-moderate, best-practice mode changes | Additional evidence unlikely to be necessary for small-to-moderate, best-practice migrations | Evaluates mode comparability within overall fit-for-purpose regulatory framework |
Condition One: What Counts as Sufficient Existing Evidence?
The first prong of the 2023 ISPOR rule requires sufficient existing evidence that the change in data collection mode has not affected the PROM's measurement properties. Sponsors cannot assume comparability from a vendor screenshot; they need a cited evidence summary for the relevant scale type and mode pair, with the limits of that literature stated.
The Meta-Analytic Foundation: Muehlhausen et al. (2015)
The public meta-analytic foundation for paper-versus-electronic PROM scores is Muehlhausen and colleagues (2015). The review covered studies conducted between 2007 and 2013 (later device ecosystems and BYOD app stores sit outside that sampling frame) and extracted 435 individual correlations and 307 standardized mean difference estimates.
Pooled agreement, with high heterogeneity: Across all 435 correlation estimates, the pooled coefficient was r = 0.88 (95% CI 0.87 to 0.88), with individual ICCs ranging from 0.65 to 0.99. Those correlations were highly heterogeneous (I2 = 93.8), so 0.88 is an average across diverse instruments rather than a universal ICC target for a new study.
Small pooled mean differences: Across 307 standardized mean differences, the pooled SMD was 0.037 (95% CI 0.031 to 0.042), about 0.04 SD, or about 1.8% of scale range (SMD I2 = 33.5, much less heterogeneous than the correlations). That is a small average difference; it is not a license to treat 0.04 SD as smaller than every minimally important difference in every indication.
Platform-stratified results are not identical: On study-averaged platform correlations (Table 3, N = 61), paper-versus-tablet/touch-screen agreement was highest at r = 0.890 (95% CI 0.876 to 0.902; pooled SMD 0.020), paper-versus-web/PC was r = 0.886 (95% CI 0.879 to 0.893; pooled SMD 0.038), paper-versus-PDA/handheld was r = 0.851 (95% CI 0.830 to 0.859) with a larger pooled SMD of 0.106, and paper-versus-IVRS was r = 0.845 (95% CI 0.824 to 0.864; pooled SMD 0.053). The authors described IVRS agreement as still in a good-agreement range (lower 95% CI at least 0.82) but slightly lower than tablet.
BYOD Cross-Device Equivalence: Byrom et al. (2018)
Muehlhausen et al. 2015 did not sample modern BYOD app ecosystems. Byrom and colleagues (2018) reported a single-site, single-visit, three-way crossover in 155 adults with chronic conditions causing daily pain or discomfort, comparing paper, a provisioned site device, and the participant's own smartphone.
The trial evaluated common response-scale types—visual analogue, numeric rating, verbal rating (including 3-category and 6-category VRS items), and Likert—not a named copyrighted instrument:
Broad Concordance Across Modes: Across all response scale types, item-level intraclass correlation coefficients (ICCs) between paper, BYOD, and provisioned site devices ranged from 0.79 to 0.98.
Paper vs. Personal Smartphone (BYOD): Paper-versus-BYOD ICCs ranged from 0.81 to 0.98, with one item (a 6-category VRS) whose 95% CI lower bound dropped below 0.75 (0.74 to 0.86).
BYOD vs. Provisioned Site Hardware: BYOD-versus-site-device ICCs ranged from 0.79 to 0.98, with one item (again a 6-category VRS) whose 95% CI lower bound was 0.72 to 0.85. Those results support measurement association for common scale types under a supervised crossover; they do not automatically cover pediatric-only populations, unsupervised field diaries, or a novel interactive scale.
Condition Two: Faithful Migration, Class 1 Standards, and Expert Screen Review
Citing published equivalence literature satisfies only the first half of the 2023 ISPOR rule. Condition Two requires documented proof of a faithful migration—demonstrating that the software implementation accurately preserves the psychometric integrity, cognitive demand, and visual clarity of the original validated instrument.
Mowlem et al. (2024): Class 1 vs. Class 2 Implementation Practices
In January 2024, the Critical Path Institute's Electronic Clinical Outcome Assessment Consortium published consolidated electronic-implementation and migration practices (Mowlem et al., Value in Health). The consortium classified each practice by whether skipping it could affect comparability, regulatory acceptability, or usability (Class 1) versus whether it mainly optimizes usability (Class 2). Class 1/Class 2 is a consortium grouping, not an FDA checklist:
| Practice cluster | Class 1 (skipping can affect comparability, regulatory acceptability, or usability) | Class 2 (mainly optimizes usability) |
|---|---|---|
| Owner rules and item content | Meet copyright/license electronic-implementation requirements; keep response-option order; keep text emphasis consistent with paper; implement only the minor wording changes needed for the new mode | Display title and copyright; keep font consistent within a screen; keep response numbering only when it encodes the response value |
| Screen containment | Make each displayed item self-contained so the question and its responses can be used together; allow enough space for translations when scrolling is disabled | Single item per screen is industry practice and is classified Class 2, not Class 1 |
| Response mechanics and navigation | Equal-sized and equally spaced hit spots; no default response; do not auto-advance; allow going back to the previous screen; on-device instructions/training | Progress indicator; consistent navigation-button placement; consistent device orientation |
| Scrolling | Not a Class 1 prohibition. Some owners forbid scrolling; some small screens need it for readable font. Mowlem reports no published evidence that vertical scrolling harms comparability | Systems should support both scrolling and non-scrolling implementations (Class 2 system functionality) |
| Completeness controls | Capture a data point for all items; prevent out-of-range or illogical responses; include a final save-and-submit screen | Avoid free-text entry where possible |
Mowlem et al. state that if at least all Class 1 practices are followed, the Consortium believes comparability studies are not necessary when migrating common response-scale types from paper to electronic format, and that where sufficient comparability evidence exists and these practices have been followed, further mode-comparability testing is not necessary, including mixed-mode and BYOD designs. That is not a reason to treat vertical scrolling as an automatic forfeit: Mowlem classifies support for scrolling and non-scrolling implementations as Class 2, notes that some small screens need scrolling for readability, and reports no published evidence that vertical scrolling negatively affects comparability—while also noting that some copyright holders forbid scrolling. Skipping actual Class 1 practices (auto-advance, default selections, unequal hit spots, or ignoring owner electronic-implementation requirements) is what reopens comparability work.
Replacing Routine Patient Testing: The Structured Expert Screen Review
Muehlhausen and colleagues (2018) synthesized all cognitive-interview and usability studies performed by one contract research organization between 2012 and 2015: 53 studies comprising 68 unique instruments and 101 instrument evaluations, spanning multiple therapeutic areas (including respiratory, gastrointestinal, oncology, rheumatology, and migraine, plus some healthy-volunteer work). The archive is not a randomized equivalence program across all vendors.
Forty-nine of 53 studies (92%) recommended no changes in display clarity, navigation, operation, or unassisted completion; reported usability findings were described as eliminable by good product design (size, location, and responsiveness of navigation buttons). Five of 53 studies (9%) identified minor cognitive-interview findings that might affect measurement properties; the authors said all could be addressed by ePRO best practice such as eliminating scrolling, ensuring appropriate font size, ensuring suitable VAS line thickness, and providing suitable instructions.
Muehlhausen et al. (2018) concluded it is possible to relax the need to routinely conduct cognitive-interview and usability studies for minor migrations when design best practice is applied. O'Donohoe et al. (2023) and McLeod and Rockwood (2023) treat a documented expert screen review by a qualified individual—against best-practice standards and, where they exist, instrument-owner requirements—as one way to show that those practices were followed. An expert screen review is documentation of migration quality, not independent proof that a specific PROM is fit for a new concept of interest, and it is not regulatory acceptance of the endpoint.
Mixing Electronic Modes and BYOD vs. The Residual Hazard of Paper Field Diaries
A foundational insight of the 2023 ISPOR framework is its sharp distinction between mixing digital interfaces versus mixing digital interfaces with paper.
Electronic-to-Electronic Mixing and BYOD Flexibility
When a protocol allows personal smartphones (BYOD), web browsers, or provisioned clinic tablets, the comparability question is whether completing the same PROM on different screen-based devices changes the scores. O'Donohoe et al. (2023) state that if faithful-migration best practices have been followed, in many instances BYOD and mixing electronic modes should not require additional comparability evidence, for both between-participant and within-participant mixing (including later provision of a site device when a personal device is lost, broken, too old, or too small). Mixing is still not described as inherently desirable.
Consequently, both between-participant mixing (e.g., 75% of trial participants using BYOD smartphones and 25% using provisioned handhelds) and within-participant mixing (for example, a later switch to a provisioned device when a personal device is lost, broken, too old, or too small) should not, in many instances, require additional comparability evidence if best practices were followed for the supported form factors.
The Paper Exception: Why Field Diaries Remain High-Risk
Eremenco et al. (2014) strongly discouraged mixing paper and electronic field-based instruments and required planned (not ad hoc) implementation plus SAP evaluation of whether mixed-mode data can be pooled. O'Donohoe et al. (2023) keep that paper-field caution as a data-quality concern. The operational problems that sit behind that caution include:
Absence of Contemporaneous Audit Trails: Electronic diaries can record contemporaneous completion times. Unsupervised paper diaries do not, which is why FDA's 2009 PRO guidance says that when an unsupervised diary is used, the Agency plans to review what steps ensure entries are made according to the trial design and not, for example, just before a clinic visit.
Skip-Pattern Failures and Ambiguous Entries: Electronic systems can prevent out-of-range or illogical responses and can enforce branching. O'Donohoe et al. explicitly flag missing or ambiguous paper entries as a data-quality risk even for site-based paper mixed with electronic collection.
Secondary Transcription Error Variance: Paper entries still have to be transcribed into the study database, which adds a transcription-error pathway that electronic capture avoids. That is an operational data-quality issue, not proof that paper and electronic scores are psychometrically incomparable when both are completed under supervision.
Supervised, site-based paper as a backup during an observed visit is the mix O'Donohoe et al. treat as potentially suitable, with the missing-or-ambiguous-entry caveat. Unsupervised home paper diaries mixed with electronic capture remain the high-risk exception the 2014 report strongly discouraged and that 2023 does not clear.
When New Empirical Testing Remains Mandatory
The 2023 ISPOR rule is conditional, not a blanket exemption. Additional testing should still be considered in at least the following situations:
Construct-Altering Modifications: Any adaptation that modifies the core clinical concept or cognitive demand of an instrument falls outside the mode comparability shortcut. In its finalized October 2025 PFDD Guidance 3, the FDA emphasized that altering an instrument's recall period (e.g., changing a daily 24-hour worst-pain assessment into a 7-day recall) does not constitute a minor format adjustment; it in effect creates an entirely new clinical outcome assessment. Similarly, altering question wording, changing response options, or modifying scoring algorithms invalidates previous validation evidence and demands de novo qualitative and psychometric research (as discussed in our guide to evaluating fit-for-purpose COAs under PFDD Guidance 3).
Novel, Complex, or Interactive Response Formats: When migrating instruments that rely on complex graphical interactions—such as interactive anatomical mannequins for localized joint pain, multi-dimensional circular dials, or gamified sliders—the absence of published meta-analytic equivalence means the 2023 shortcut does not apply. Cognitive interviewing, usability testing, and, where needed, quantitative comparability work should be considered.
Violations of Class 1 Faithful Migration Standards: If implementation skips Class 1 practices—auto-advance, a default response, unequal hit spots, ignored owner electronic-implementation requirements, or splitting a self-contained item so that the question and its responses cannot be used together—the 2023 shortcut does not apply. Cognitive interviewing, usability testing, or quantitative comparability work should then be considered, depending on the measure and intended use. Vertical scrolling by itself is not that trigger in Mowlem et al. 2024.
Vulnerable or Cognitively Impaired Populations: Meta-analytic and BYOD crossover evidence is strongest in supervised adult samples. Pediatric-only, cognitively impaired, or other populations that were not covered by the cited studies still warrant dedicated usability and, where the 2023 evidence condition fails, comparability work.
Registrational Endpoints with Significant Mode Mixing: When modes will be mixed, Eremenco et al. 2014 still required the SAP to address mixing and to evaluate whether data can be pooled before evaluating treatment effects. FDA's 2009 PRO guidance remains listed current and says the Agency intends to review whether treatment effect varies by method or mode when multiple methods or modes are used in one trial. PFDD Guidance 3's 'unlikely to be necessary' language does not cancel that review interest.
How FDA Interprets Mode Comparability: Reconciling 2009 and 2025 Guidance
FDA guidances are nonbinding recommendations unless a cited statute or regulation applies. The still-listed 2009 PRO guidance and the October 2025 final PFDD Guidance 3 have to be read together; neither is a labeling guarantee for a mixed-mode PROM endpoint.
The 2009 PRO Guidance: Reviewing Treatment Effects Across Modes
FDA's December 2009 guidance, Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims, remains listed current. Under content validity (section III.D.2, Data Collection Method and Instrument Administration Mode), the Agency distinguishes administration mode (self, interview, or both) from data-collection method (paper-based, computer-assisted, telephone-based). It states that it intends to review comparability when multiple methods or modes are used in one trial to determine whether the treatment effect varies by method or mode. Separately, the guidance lists changing an instrument from paper to electronic format among modifications that can alter how patients respond, while also stating that not every small change in application or format requires extensive new measurement-property studies and that additional qualitative work may be adequate depending on the modification. Sponsors should still be able to provide the exact instrument version, including screenshots of electronic PRO instruments where relevant.
October 2025 Final PFDD Guidance 3: Mode-comparability paragraph
In October 2025, the FDA released its finalized Patient-Focused Drug Development Guidance 3: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments uses O'Donohoe et al. 2023 for its mode-of-administration comparability paragraph. That paragraph is in scope here; the rest of Guidance 3 is the fit-for-purpose selection/modification job covered in a separate published article:
Direct Recognition of Empirical Comparability: The guidance states that empirical studies across a range of settings generally support comparability of measurement properties across paper-based, tablet computer, or a patient's own device, citing O'Donohoe et al. 2023.
Unlikely-to-be-necessary statement: Guidance 3 states that additional supportive evidence for comparability across modes is unlikely to be necessary if (a) the changes are small to moderate (examples: item format or presentation, or paper-based to interactive voice response) and (b) sponsors followed best practices for migrating measures to different assessment platforms, citing among others Eremenco et al. 2014, Mowlem et al. 2024, Byrom et al. 2019, and C-Path ePRO Consortium materials.
Clear Boundary on Substantial Changes: The guidance carefully distinguishes benign format adaptations (e.g., moving from multiple paper items per page to a single item per tablet screen) from substantive modifications (e.g., shifting recall periods from 1 day to 7 days), explicitly noting that the latter may create an entirely distinct measurement tool.
Pre-Specifying Mixed Modes in the Statistical Analysis Plan (SAP)
When a clinical trial incorporates mixed data collection modes—whether deliberately (for example, BYOD with a provisioned backup) or operationally (for example, in-clinic tablet backup during technical failure)—the statistical analysis plan should state, before database lock, how mixed scores will be attributed, described, pooled, and probed for mode effects.
ICH E9(R1) is about estimands and same-estimand sensitivity analysis, not a mixed-mode checklist. If modes will be mixed, the SAP still needs planned pooling rules. The following specifications are an operator application of that requirement and of FDA's 2009 interest in treatment effect by mode—they are not numeric thresholds invented from ICH E9(R1):
Granular Mode Attribution in Data Specifications: Record the collection method actually used (BYOD app, provisioned device, web, supervised paper backup) as metadata for each assessment, so pooling and mode-effect analyses are reconstructable.
Descriptive Balance Diagnostics: The SAP should specify baseline and on-treatment summaries of collection method by randomized treatment arm, study center, and region. BYOD is often self-selected, so imbalance is possible; the job is to describe that distribution and assess confounding, not to assume mode was randomized.
Evaluating Treatment-by-Mode Interactions: For primary and key secondary endpoints, complement the primary model with a pre-specified look at treatment effect by collection method. A non-significant interaction is not proof of score invariance, but it is how teams answer the review question FDA described in 2009. No universal ICC, alpha, or sample-size threshold for that analysis is stated in the cited texts.
Pre-Specified Sensitivity Rules for Emergency Paper Backups: If supervised in-clinic paper backups are used, pre-specify a sensitivity analysis that excludes paper-derived assessments so reviewers can see whether primary conclusions depend on those records.
The Reconstructable Evidence Dossier: A 5-Step Operational Protocol
Before first mixed-mode collection, teams should be able to reconstruct why new comparability testing was or was not done. A vendor screenshot, a provisioned-versus-BYOD logistics memo, a computerized-system validation file, or a DHT verification package is not that record.
graph TD
A["Step 1: Classify Migration Scope"] --> B{"Construct Altered?<br>(Recall, Wording, Scoring)"}
B -- Yes --> C["De novo qualitative and quantitative work"]
B -- No --> D["Step 2: Compile Scale Evidence Base<br>(Muehlhausen 2015, Byrom 2018)"]
D --> E{"Common Scale Type?<br>(VAS, NRS, Likert, VRS)"}
E -- No --> F["Conduct Scale-Specific<br>Comparability Study"]
E -- Yes --> G["Step 3: Conduct Structured Screen Review<br>(Class 1 Rules & Owner Specs)"]
G --> H{"Class 1 Adherence Confirmed?"}
H -- No --> I["Perform Focused Usability /<br>Cognitive Debriefing"]
H -- Yes --> J["Step 4: Execute Testing Decision Gate<br>(O'Donohoe 2023 & PFDD 3 Rationale)"]
J --> K["Step 5: Align Protocol & SAP<br>(Mode Metadata, Pooling & Sensitivity)"]
K --> L["File Complete Reconstructable Dossier in TMF"]Step-by-Step Implementation Protocol
Step 1 — Classify the Migration Scope: Formally document whether the proposed adaptation represents a minor format adjustment (screen sizing, typography, pagination), moderate layout adjustment, or substantial construct modification (altered recall period, reworded items, changed scoring). If construct-altering, route immediately to de novo validation.
Step 2 — Compile the Scale Evidence Base: Retrieve and catalog peer-reviewed literature demonstrating measurement comparability for the target response scales (e.g., Muehlhausen et al. 2015 for general electronic concordance; Byrom et al. 2018 for BYOD cross-device equivalence).
Step 3 — Conduct and Document the Structured Expert Screen Review: Commission an outcome assessment specialist to conduct a screen-by-screen audit of the built eCOA application across all supported device viewports, verifying adherence to instrument-owner guidelines and Mowlem et al. (2024) Class 1 rules.
Step 4 — Execute the Testing Decision Gate: If Class 1 practices were followed and existing evidence covers the scale type and mode pair, document the rationale for not commissioning a new comparability study, citing O'Donohoe et al. 2023 and the PFDD Guidance 3 mode paragraph as current thinking. If those conditions fail, commission cognitive, usability, or quantitative work as needed. That memo is not Agency acceptance of the endpoint.
Step 5 — Align Protocol and Statistical Analysis Plan: Specify approved collection methods in the protocol, keep unsupervised home paper diaries out of a mixed electronic plan unless a new evidence package supports them, and pre-specify pooling and mode-effect evaluation in the SAP.
