A drug failed by unreliable measurement is recorded as a drug that did not work.
A Clinical Outcome Assessment is the instrument that produces it — a structured measure of how a participant feels, functions or survives. COAs define a trial's endpoints: the evidence a regulator uses to decide whether a treatment works, and the basis of the label it is sold under.
Everything upstream — the molecule, the sites, the participants, the years — resolves to whether that one measured number separates from control. It is the most consequential measurement in the industry, and the least industrialised.

The measurement breaks in five places.
Each break has a known cause and a known fix. Our solutions are built around them.
The instrument
Instrument design constrains everything downstream. Most COAs in use were built for paper and are licensed by a handful of rights-holders, so the people running the trial cannot adapt them for digital or decentralised capture. The SGRQ, the standard COPD questionnaire, derives from a 1991 paper instrument; when it was digitised in 2015 the stated validation goal was that participants respond to it in the same way as they had on paper — equivalence to the page, not fitness for the screen (Jones et al., Respir Med 1991).
eCOA development & licensing →Administration
How an instrument is administered is part of the measurement. Vary the procedure and you vary the measure — and experienced raters are routinely trained differently on the same administration rule from one protocol to the next. The ADAS-Cog, a single instrument, carries three different sets of instructions for how to rate an erroneous response on the Commands task (Connor & Sabbagh, J Alzheimers Dis 2008).
Rater training & certification →Scoring
Rater reliability drives sample size. Every point of unreliability is variance that has to be bought back with participants — raters who have never been calibrated diverge widely, and the difference is paid for in participants. Indexed to a trial of 100 participants at an intraclass correlation of 0.93, uncalibrated raters required 69 more participants to reach the same statistical power (Kobak et al., J Clin Psychopharmacol 2009).
Endpoint quality surveillance →Who gets enrolled
Unreliable baseline scoring selects the wrong participants, and participants who are not sick enough get better on placebo. The baseline score that decides eligibility is made by the site, under pressure to enrol — gatekeeper and recruiter are the same person. In one study, blinded central raters would have excluded 63 of the 122 participants the sites randomised, and measured half the placebo response (Williams et al., J Clin Psychopharmacol 2015).
Endpoint quality surveillance →Clinical adjudication
Adjudication is still a manual, cross-platform exercise. The evidence for a single endpoint event sits in ePRO, EDC, labs, imaging and safety systems that do not talk to each other, so determinations are slow, inconsistently evidenced and hard to audit after the fact. In a 10,948-participant trial where a committee screened every participant's data, it found 816 endpoint infarctions the sites had never reported — 58% of every endpoint infarction in the trial — and rejected 167 the sites had reported. Six-month mortality among the 816 was 12.6%, against 3.8% where site and committee agreed there was no event (Mahaffey et al., Curr Control Trials Cardiovasc Med 2001).
Clinical endpoint adjudication →Which of the five is your endpoint exposed to?
Bring us the protocol and the instruments. We will map the exposure and tell you what is worth covering.
Talk to us about your study