COA Instruments

PRO, ClinRO, ObsRO, or PerfO: Choosing the COA Type for the Concept of Interest

A methodological decision framework for clinical trialists to select among PRO, ClinRO, ObsRO, and PerfO measures based on the concept of interest and context of use before naming an instrument.

· · 20 min read

Editorial research-journal still life of four unbranded clinical outcome assessment objects arranged on warm parchment and oxblood linen: a blank cardstock sheet, a clinician clipboard, an observation notebook, and a mechanical stopwatch beside a wooden task block

The decision is which COA type can measure the concept before an instrument is named

When architecting clinical trial endpoints, the first outcome-measurement decision is not which named scale, proprietary questionnaire, or commercial battery to license. It is determining which of the four Clinical Outcome Assessment (COA) types—Patient-Reported Outcome (PRO), Clinician-Reported Outcome (ClinRO), Observer-Reported Outcome (ObsRO), or Performance Outcome (PerfO)—possesses the reporter authority, observational access, and methodological capability to measure the specified concept of interest in the target patient cohort. As stated in the U.S. Food and Drug Administration (FDA) COA Qualification Program Frequently Asked Questions, the type of COA to select or develop depends on the target concept of interest and the context of use, particularly the patient population. If pain intensity is the concept of interest in a patient population capable of self-reporting, a PRO is most appropriate. If clinical judgment is required to interpret an observation, a ClinRO is chosen. If the concept can only be adequately captured through daily-life observation outside a healthcare setting and the patient cannot self-report, an ObsRO is chosen. When it would be useful to observe an actual demonstration of a defined task demonstrating functional performance, a PerfO may be appropriate.

The final FDA Patient-Focused Drug Development (PFDD) Guidance 3 (October 2025), Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments, places this choice in Section III.C.1 as the first step under selecting or developing an outcome measure: sponsors and measure developers should consider which type is most appropriate for the concept of interest in the context of use. In Section IV, Table 2, the first component (Component A) of the evidence-based fit-for-purpose rationale asks sponsors to state: "The concept of interest should be assessed by [COA type] because . . ." That sentence is necessary and not sufficient. Choosing a COA type does not establish that any named questionnaire or test battery is fit for purpose, does not qualify an instrument under FDA's Drug Development Tool (DDT) qualification program, and does not constitute regulatory acceptance of an endpoint. The Qualification Program FAQ, the general COA FAQ, and PFDD Guidance 3 are listed recommendations, not regulations.

This guide provides clinical development leaders, biostatisticians, and outcome-assessment specialists with an operational decision framework for that initial type-selection fork. To preserve methodological focus, this decision rule fences off subsequent and orthogonal measurement questions addressed elsewhere across the Trial Endpoints platform:

State the meaningful aspect of health, the concept of interest, and whether self-report is feasible

A frequent flaw in clinical trial protocols is selecting a published outcome instrument simply because it was utilized in prior trials or recommended by a clinical advisory board, without formally decomposing the measurement construct. FDA PFDD Guidance 3 and the FDA Roadmap to Patient-Focused Outcome Measurement require trialists to define three sequential elements before deciding on the COA type:

  • Meaningful Aspect of Health (MAH): The specific aspect of health, disease impact, or bodily function that patients experience and care about in their daily lives. Examples include physical mobility, sustained restorative sleep, emotional equilibrium, pain burden, or gastrointestinal stability. While eliciting what matters to patients is the domain of qualitative interview research under PFDD Guidance 2, the MAH is the feeling-or-functioning aspect a COA-based endpoint is meant to help understand.

  • Concept of Interest (COI): The discrete, measurable facet of the meaningful aspect of health that the clinical trial intends to assess and evaluate for treatment effect. For example, within the meaningful aspect of physical mobility, potential concepts of interest include maximum walking distance, stair-climbing velocity, transfer independence, or perceived physical exhaustion after exertion.

  • Context of Use (COU): The clinical trial parameters governing where and how the measurement occurs, including the target disease indication, clinical staging, patient demographic profile, setting of assessment (clinic versus free-living home environment), and fundamentally, whether the target patient cohort has the developmental, cognitive, and physical capacity to provide a reliable self-report.

The determination of whether the patient can reliably self-report serves as the primary operational fork in the entire endpoint selection hierarchy. When patients possess the communicative and cognitive maturity to report their own condition, the patient is the only epistemically valid reporter for subjective sensations, unobservable feelings, and internal symptoms. Conversely, when developmental age (such as in neonates or non-verbal infants), severe cognitive impairment (such as in advanced dementia or severe encephalopathy), or acute communicative incapacity makes self-report unfeasible, trialists cannot substitute a caregiver proxy guess of internal feelings. Instead, the concept of interest itself should be recalibrated to focus on observable clinical signs, overt daily behaviors, or standardized functional performance.

The four-type taxonomy and the whose-judgment test

The FDA general Clinical Outcome Assessment FAQ (content current as of 2 December 2020) and the FDA-NIH Biomarkers, EndpointS, and other Tools (BEST) Resource glossary, restated on FDA's PFDD glossary page (current as of 5 March 2026), define a Clinical Outcome Assessment as a measure that describes or reflects how a patient feels, functions, or survives. Assessment of a clinical outcome can be made through report by a clinician, a patient, a non-clinician observer, or through a performance-based assessment. These four categories are the current public FDA and BEST taxonomy for clinical outcome measurement.

In their foundational ISPOR task-force report, Walton and colleagues (Value in Health, 2015) articulated the core methodological principle that distinguishes these four classes: whether judgment can influence the measurement, and whose judgment is exercised. Understanding whose judgment is embedded within the score clarifies which reporter class can generate the measurement:

  • Patient's Subjective Judgment (PRO): The measurement comes directly from the patient without amendment, interpretation, or clinical filtering by anyone else. It captures the patient's personal perception of their internal physical state, symptoms, and functional impacts.

  • Clinician's Professional Judgment (ClinRO): The measurement is reported by a trained healthcare professional whose specialized medical education, diagnostic judgment, and clinical interpretation are required to evaluate observable signs, physical manifestations, or psychiatric behaviors.

  • Observer's Naturalistic Observation Without Medical Judgment (ObsRO): The measurement is reported by a non-clinician observer (such as a parent, spouse, or professional caregiver) who routinely observes the patient in daily life outside a healthcare facility. Crucially, the observer records concrete, observable behaviors and events without medical interpretation or inferential guessing of internal states.

  • Standardized Task Execution Without Subjective Rating (PerfO): The measurement is based on a standardized task actively executed by the patient according to precise instructions. The score is quantified through objective parameters (such as elapsed time, distance traversed, accuracy percentage, or weight supported) rather than through rater interpretation.

It is equally critical to understand what does not constitute a COA type under this type-selection tree. Biomarkers (such as serum protein assays, blood pressure readings, histological staining, or MRI volumetrics) measure biological or physiological processes, not clinical outcomes of how a patient feels or functions. Digital health technology (DHT) sensor metrics (such as passive 24-hour actigraphy or other mobile-sensor streams) are not a fifth COA type; PFDD Guidance 3 Section II.A directs that case to FDA's December 2023 DHT guidance. Survival remains inside the BEST and general COA FAQ definition of how a patient feels, functions, or survives, but PFDD Guidance 3 footnote 18 limits that guidance to feel-or-function assessments, not how patients survive, so vital-status adjudication is outside this four-type tree.

COA TypePrimary ReporterNature of JudgmentTarget Concept ClassesRegulatory & Methodological Limits
PRO (Patient-Reported Outcome)Patient exclusively (unassisted or with non-interpretive recording accommodation)Subjective perception and self-appraisal of internal health statusSymptoms known only to the patient (pain, nausea, fatigue, itch, mood), perceived daily functioning, treatment burdenRequires capacity for reliable self-report. FDA discourages proxy reports of internal states.
ClinRO (Clinician-Reported Outcome)Trained, qualified healthcare professionalProfessional clinical judgment, medical staging, and diagnostic interpretationObservable physical signs (lesion morphology, spasticity), diagnostic symptom clusters, psychiatric affect, cognitive statusCannot directly assess internal symptoms known only to the patient. A named ClinRO still needs later fit-for-purpose evidence; rater training is a separate operational job.
ObsRO (Observer-Reported Outcome)Non-clinician observer (parent, spouse, home caregiver)Report of concrete observation in daily life; medical judgment is strictly excludedObservable daily behaviors, seizure frequency, vomiting episodes, infant crying duration, assistance required for daily tasksParticularly useful when the patient cannot self-report. Observer reports observable signs, events, or behaviors, not inferred internal feelings.
PerfO (Performance Outcome)Patient executing standardized task; scored by trained proctor or automated systemStandardized task metric; subjective rater judgment is minimized or eliminatedDemonstrated functional capacity (for example, distance walked in six minutes, visual acuity, or word-list recall)Patient needs to be able to follow the instructions. Evaluates demonstrated capability in a structured setting, not free-living daily habits.
graph TD
    Start["Define Meaningful Aspect of Health & Concept of Interest"] --> Q1{"Is the concept an unobservable internal state or subjective symptom?"}
    
    Q1 -- "Yes (Pain, Fatigue, Nausea, Itch)" --> Q2{"Can the patient cohort reliably self-report?"}
    Q2 -- "Yes" --> PRO["Select PRO (Patient-Reported Outcome)<br/>Direct patient report; proxy inference is not a PRO"]
    Q2 -- "No (Infants, Severe Cognitive Impairment)" --> Refocus["Reframe Concept of Interest:<br/>Shift from internal sensation to observable signs or behaviors"]
    
    Q1 -- "No (Observable Sign, Behavior, or Functional Ability)" --> Q3{"Does evaluating the observation require professional medical judgment?"}
    
    Refocus --> Q3
    
    Q3 -- "Yes (Pathology Staging, Spasticity, Psychiatric Affect)" --> ClinRO["Select ClinRO (Clinician-Reported Outcome)<br/>Qualified healthcare professional applies clinical interpretation"]
    Q3 -- "No (Direct Observation or Standardized Task)" --> Q4{"Does the concept evaluate demonstrated capability during a standardized task?"}
    
    Q4 -- "Yes (Walking Speed, Memory Recall, Manual Dexterity)" --> PerfO["Select PerfO (Performance Outcome)<br/>Patient executes standardized task under uniform instructions"]
    Q4 -- "No (Daily-Life Behavior Outside Healthcare Setting)" --> Q5{"Can the patient report the behavior themselves?"}
    
    Q5 -- "Yes" --> PRO
    Q5 -- "No (Patient unable to report)" --> ObsRO["Select ObsRO (Observer-Reported Outcome)<br/>Caregiver records concrete observable events without medical judgment"]
Canonical decision hierarchy for selecting a Clinical Outcome Assessment type based on concept of interest and context of use.

PRO when the concept is known only to the patient and self-report is reliable

Under the FDA-NIH BEST Resource and the FDA 2009 PRO Guidance, a Patient-Reported Outcome is defined as a measurement based on a report that comes directly from the patient about the status of a patient's health condition without amendment or interpretation of the patient's response by a clinician or anyone else. FDA's guidance explicitly establishes that symptoms—defined as subjective evidence of disease that can be noticed and known only by the patient—can only be measured by PRO instruments. When an articulate adult patient suffers from pain, chronic fatigue, nausea, depressive mood, headache, or neuropathic dysesthesia, no clinician rating, laboratory biomarker, or caregiver observation can substitute for the patient's own evaluation.

PFDD Guidance 3 Appendix A reinforces that a PRO is the optimal COA type for capturing feelings known only to the patient, the patient's personal perspective on daily functioning, or the perceived degree of treatment burden, provided that the target cohort can reliably self-report. Importantly, regulatory guidance recognizes that certain clinical operations require administrative flexibility without violating PRO integrity:

  • Interviewer Administration: An interviewer may administer a PRO survey verbally (for example, via telephone or during an in-clinic visit), provided that the interviewer is trained to read questions and response options verbatim and records strictly the patient's unprompted response without rephrasing, clarifying, or interpreting the answer.

  • Disability and Accessibility Accommodations: Footnote 15 of PFDD Guidance 3 clarifies that assistance necessary to accommodate physical disability—such as an aide reading questions to a visually impaired participant, holding a tablet for an arthritic patient, or recording keystrokes dictated by a patient with severe tremor—remains a valid PRO as long as the assistant records the patient's exact response without exercising independent judgment.

In contrast, FDA guidance maintains a firm fence against "proxy-reported outcomes". The 2009 PRO guidance defines a proxy report as a report by someone other than the patient who attempts to report as if he or she is the patient. Caregivers, parents, and healthcare personnel cannot experience another individual's internal sensory or emotional state. FDA discourages proxy-reported outcome measures particularly for symptoms that can be known only by the patient. When patients cannot self-report, PFDD Guidance 3 recommends an ObsRO of observable behavior rather than a proxy report of the patient's inner experience.

ClinRO when clinical judgment is required to interpret an observation

A Clinician-Reported Outcome is defined by the BEST Resource and PFDD Guidance 3 Appendix C as a measurement based on a report that comes from a trained healthcare professional after observation of a patient's health condition. Most ClinRO measures involve a clinical judgment or interpretation of observable signs, physical behaviors, or diagnostic manifestations related to a disease or condition. The defining attribute of a ClinRO is the necessity of specialized professional clinical training: the rater must possess medical or psychological expertise to synthesize clinical observations into a reliable, standardized rating.

Typical applications where a ClinRO is the indicated type include assessing morphological tissue damage (such as the extent of psoriasis plaques), evaluating neurological motor tone (such as grading spasticity), rating clinical signs that require professional interpretation, or completing a clinician global assessment of current status. In each case, an untrained layperson or the patient themselves cannot supply that specialized clinical judgment. Named scales mentioned in BEST or PFDD Guidance 3 appendices illustrate the type; they are not treated here as fit-for-purpose or qualified.

However, regulatory guidance and the ISPOR ClinRO task-force report (Powers et al., Value in Health, 2017) emphasize a methodological boundary: ClinRO measures cannot directly assess symptoms that are known only to the patient. A clinician cannot know exactly how patients feel or experience symptoms. Substituting a physician's rating for an articulate patient's self-reported symptom is a construct mismatch, not a ClinRO substitute for a PRO.

Powers and colleagues provide a benchmark illustration of how COA type selection varies across patient populations using the example of pruritus (itch) in atopic dermatitis:

  • In adolescent and adult populations capable of self-reporting, itch intensity is a subjective internal sensation that should be evaluated using a PRO.

  • In infants or young children who cannot self-report, itch intensity cannot be measured directly. Powers notes that a ClinRO or ObsRO is then a better type for clinical signs or behaviors; PFDD Guidance 3 similarly pairs a caregiver ObsRO of scratching behavior with a PRO of itch intensity only for the subset who can validly self-report.

Finally, instruments cited in BEST or FDA guidance appendices (such as the Psoriasis Area and Severity Index [PASI] or the Hamilton Depression Rating Scale [HAM-D]) serve solely as illustrative examples of the ClinRO class. Their appearance in glossaries does not constitute qualified status, does not imply blanket FDA endorsement, and does not establish that the instrument is fit for purpose in a specific context of use.

ObsRO for observable daily-life signs when the patient cannot self-report—not a proxy

Under BEST and PFDD Guidance 3 Appendix B, an Observer-Reported Outcome is defined as a measurement based on a report of observable signs, events, or behaviors related to a patient's health condition by someone other than the patient or a healthcare professional. ObsRO measures are completed by individuals who interact with the patient in their natural daily environment, such as parents, spouses, partners, or home caregivers. According to FDA's COA Qualification Program FAQ, an ObsRO is chosen under two conditions:

  • The concept of interest can only be adequately captured by observation in daily life outside a healthcare setting; and

  • The patient cannot report for himself or herself due to developmental age, cognitive impairment, or severe communicative disability.

The cornerstone of valid ObsRO measurement is the separation between concrete behavioral observation and inferential proxy guessing. An ObsRO should not require the observer to infer or interpret the patient's internal psychological or sensory state. Valid ObsRO items ask observers to report observable signs, events, or behaviors:

  • Valid ObsRO Concepts: Counting the number of overt convulsive seizures in a 24-hour log; recording episodes of postprandial emesis; timing the duration of continuous crying spells in an infant; noting whether a toddler voluntarily consumes solid food; or logging whether an elderly dementia patient required physical assistance to put on their clothing.

  • Invalid Proxy Inferences: Asking a mother to rate "how intense the child's abdominal pain was" on a 1-to-5 scale; asking a spouse to grade "how anxious the patient felt during the morning"; or asking a caregiver whether a non-verbal patient experienced nausea. Observers cannot observe pain, anxiety, or nausea; they can only observe grimacing, restlessness, or vomiting.

PFDD Guidance 3 Appendix B encapsulates this rule: itch intensity is known only to the patient and should be assessed using a PRO. FDA strongly discourages a proxy inference of itch intensity. A caregiver may instead report observable scratching behaviors as an ObsRO. The same guidance's pediatric example pairs a caregiver ObsRO of scratching behavior with a PRO of itch intensity only for the subset of patients who can validly self-report.

Furthermore, those sources do not establish arbitrary numeric age cutoffs (such as declaring that all children under age 7 cannot self-report while all children over age 8 can) or universal cognitive thresholds. Whether self-report is feasible is a context-of-use judgment. PFDD Guidance 3 Appendix B recommends considering whether observers have sufficient contact with the patient, and it distinguishes observable-behavior items from proxy inferences of inner experience; it does not prescribe a numeric age cutoff or a universal cognitive score.

PerfO when the concept is functional performance on a standardized task

The BEST Resource and PFDD Guidance 3 Appendix D define a Performance Outcome as a measurement based on a standardized task or tasks actively undertaken by a patient according to a set of standardized instructions. Unlike PROs or ObsROs that capture free-living daily experiences across variable recall periods, a PerfO evaluates what a patient is physically, cognitively, or sensory capable of demonstrating under controlled, reproducible conditions.

In their comprehensive ISPOR task-force report, Edgar and colleagues (Value in Health, 2023) established recommendations on determining when a PerfO is the optimal COA type. A PerfO is indicated when the concept of interest is functional capacity that is best demonstrated through standardized task execution, particularly in circumstances where:

  • Objective Demonstration Outweighs Subjective Recall: Patients may have difficulty recalling functioning over a recall period; PFDD Guidance 3 Appendix D notes that real-time PerfO tasks are not vulnerable to that recall error.

  • Standardized Task Versus Free-Living Heterogeneity: Daily-life observation varies with setting and assistance. Appendix D notes that standardized tasks in a controlled environment may be less influenced by that heterogeneity than free-living reports.

  • Patient Has Limited Capacity for Self-Report but Can Follow Instructions: In pediatric cohorts or patients with mild-to-moderate cognitive impairment who cannot complete complex written questionnaires, patients can often still follow simple, direct instructions to execute a physical or cognitive task.

Typical concept classes where a PerfO may be considered, per PFDD Guidance 3 Appendix D and the BEST/PFDD glossary examples, include mobility, memory, and visual acuity—illustrated by distance walked in six minutes, a timed walk, word-recall or digit-symbol tasks, and visual-acuity testing. Those named tasks illustrate type; they are not treated here as qualified or fit-for-purpose batteries.

Methodologically, outcome specialists should distinguish PerfO measures from both ClinRO ratings and digital health technology (DHT) sensors:

  • PerfO versus ClinRO: In a PerfO assessment, a trained individual may administer the task and record the outcome, but the investigator does not apply judgment to quantifying the performance (Walton 2015). Appendix D's examples include meters walked in six minutes or words recalled. In contrast, a ClinRO requires the clinician to apply medical judgment to rate what is observed, including some standardized tasks whose scores still depend on clinical interpretation.

  • PerfO versus Passive DHT Sensors: A PerfO requires an active, episodic task undertaken in response to standardized instructions (e.g., walking as briskly as possible for six minutes). A passive wearable sensor (such as a wrist-worn accelerometer collecting continuous movement data over seven days during uninstructed daily living) captures free-living behavior, not a standardized task. Sensor-derived metrics fall under FDA's digital health technologies framework, not the PerfO classification.

A PerfO can be considered when the patient is able to follow the instructions to perform the required task. Appendix D notes that if cognitive ability may interfere with task performance, sponsors should consider whether the selected PerfO is fit-for-purpose. Locking a standardized task protocol does not constitute FDA qualification, device clearance, or proof that the battery is fit for purpose.

What type choice does not change, and what it does not claim

Selecting the appropriate COA type based on the concept of interest and context of use is a necessary first step in protocol design, but trial teams should understand what this decision does and does not accomplish. PFDD Guidance 3 and the listed FDA FAQs set the following boundaries:

  • Electronic Mode Does Not Change COA Type: Transitioning a questionnaire or scale from paper to a smartphone app, web portal, or tablet (eCOA) changes the mode of administration, not the COA type. As PFDD Guidance 3 Section II.A states, the measurement source remains the patient, clinician, observer, or task. Operational questions regarding screen size, layout comparability, and electronic audit trails belong to mode-comparability evidence, not to a change of reporter class.

  • Disability Assistance Does Not Compromise Type: Providing transcription, verbal reading, or physical device handling to accommodate motor or sensory disability preserves PRO status, provided the assistant records the patient's dictated responses verbatim without editorial interpretation.

  • DHT Sensors Are Not a Fifth COA Type: Continuous digital sensors (such as a mobile sensor or wrist-worn accelerometer) provide digital health technology measurements. They do not constitute a fifth COA type, and PFDD Guidance 3 directs that case to FDA's December 2023 DHT guidance.

  • Multi-Component Endpoints Require Independent Type Rationales: PFDD Guidance 3 states that combining scores from several measurement types into a single multi-component endpoint is beyond that guidance's scope. Each underlying component still needs its own type rationale; combination rules are not solved by picking one type.

  • Component A Is Necessary but Far From Sufficient: Formulating the type-selection rationale under Table 2 Component A of PFDD Guidance 3 ("The concept of interest should be assessed by [COA type] because . . .") fulfills only the first of eight Table 2 components comprising an evidence-based rationale. Components B through H—coverage of the concept, appropriate administration, respondent comprehension, scoring, construct-irrelevant influence, measurement error, and correspondence of scores to the related meaningful aspect of health—should still be considered for the named instrument. Guidance 3 states those components are likely but not necessarily needed in every rationale.

  • Type Choice Is Not Regulatory Acceptance: Selecting PRO, ClinRO, ObsRO, or PerfO does not establish that a named instrument is validated, does not qualify an assessment tool under FDA's COA Qualification Program, and does not guarantee that regulatory reviewers will accept the resulting endpoint to support product approval or labeling claims.

With the COA type chosen from the concept of interest and context of use, clinical development teams can search existing measures, evaluate published evidence, or consider modification under the remaining PFDD Guidance 3 rationale components, detailed in Fit-for-Purpose COAs: Using FDA PFDD Guidance 3 to Select or Modify an Instrument.