Tech blog
We score our clinical AI by what clinicians correct
Contents
- The vocabulary, first
- Passing test recordings is not evidence of accuracy
- Rank dates by strength of evidence, not by a fixed precedence
- Values derived from a calendar belong in code
- Nine rounds, all 56 recordings, scored by machine
- After release, score on what clinicians corrected
- The practices, condensed
- Summary
この記事は日本語でもお読みいただけます。
Ask a clinical AI vendor how they verify accuracy and the easiest answer is a count: how many test recordings the system passes. It is measurable, and it grows reassuringly.
The count is not evidence. Failures in shapes your test set does not contain stay invisible no matter how many recordings you add. We build clinical documentation from consultation audio, and rebuilding how the next-appointment date is chosen forced us to measure that gap directly.
The vocabulary, first
These terms carry the rest of the post.
| Term | What it means |
|---|---|
| Transcription | Turning audio into text. The same audio does not yield identical text on every pass |
| Evaluation set | The inputs and expected answers used to measure accuracy. A skewed set measures only its own skew |
| Sampling bias | The collected audio not matching the distribution of real consultations. Adding volume does not fix it |
| Regression | Something unrelated to the fix getting worse than before |
| Mutation testing | Deliberately breaking a guard and confirming the check fails |
| Denominator | The base of a rate. When it moves, the rate moves with it for no real reason |
| History table | One row per saved version of a record. Adjacent rows reveal what a human changed |
| Deterministic substitution | Rewriting part of generated text by fixed rules, so the same input always yields the same output |
| Read-only role | A database connection that cannot write, so the measuring tool cannot damage what it measures |
Passing test recordings is not evidence of accuracy
The next-appointment date is copied verbatim onto what the patient takes home. After improving it once, we kept counting production records — specifically, how often a clinician edited the next-appointment line by hand.
The correction rate had gone from 4.8% to 9.6%. Our test set had been passing 16 to 18 of 20 recordings just before release. Had we only watched that number, we would have called it an improvement.
The reason is not subtle. Those 20 recordings contained the failure shapes we knew about. The 157 production consultations contained something else: appointments that get revised mid-conversation.
“Come back in four weeks.” “…actually we’re closed that day, let’s say October 19th.”
Take the first phrase and you get October 12th. What was actually agreed is October 19th. If that shape is absent from your test set, no number of passing recordings will reveal it.
Moving the measurement into production pulls those cases in automatically. A record a clinician corrected contains both answers: what the model wrote, and what the clinician replaced it with. Nobody has to author ground truth.
Rank dates by strength of evidence, not by a fixed precedence
Handling revised appointments meant changing how a date is chosen at all.
The previous design ranked relative expressions (“in four weeks”) above explicitly stated dates (“October 19th”). Speech recognition confuses month names — in Japanese, September and May are a single vowel apart — whereas a relative expression survives that confusion. The reasoning was sound.
But that ordering discards whatever was agreed at the end of the conversation. So we reordered the candidates by how strong the evidence is, rather than by which form is generally more robust.
| Rank | Signal | Example | Why here |
|---|---|---|---|
| 1 | Stated date whose weekday matches the calendar | “Sunday, October 19th” | Date and weekday agreeing by accident is unlikely. Agreement rules out mishearing |
| 2 | Date settled at the end of the exchange | “Let’s make it the 19th after all” | The conversation itself is the evidence |
| 3 | Stated date within 7 days of a directional relative | “In four weeks… on October 19th” | Proximity indicates both refer to one appointment |
| 4 | Month corrected via weekday | “Thursday the 10th” | Usable when only one month within two months fits |
| 5 | Stated date disambiguated by visit context | “Next visit is October 1st” | Separates the appointment from other dates in the same talk |
| 6 | Relative expression alone | “In four weeks” | Used only when no stated date exists |
Month-scale expressions appear nowhere in that table. “In a month” does not mean the same date next month; it means somewhere in the next month. When the clinician speaks in months, we keep the phrasing and generate no date. In production, roughly 30% of month-scale lines that carried a computed date were corrected by hand, while no one ever added a date to a month-scale line that had none.
That ranking is also the shape of the implementation. Rules are applied top-down and the first one that matches wins.
// Ordered by strength of evidence. The order *is* the specification,
// so adding a rule means reviewing where in the order it belongs.
const RULES = [
{ id: 'explicit_with_weekday', note: 'date and weekday agree with the calendar' },
{ id: 'agreed_last', note: 'the date settled at the end of the exchange' },
{ id: 'explicit_near_relative', note: 'stated date within 7 days of a relative one' },
{ id: 'weekday_corrected', note: 'weekday disambiguates a misheard month' },
{ id: 'explicit_unique', note: 'exactly one candidate in an appointment context' },
{ id: 'relative_only', note: 'used only when no stated date exists' },
] as const; // matching logic omitted
// Keep which rule fired, not just the date it produced
const decide = (u: Utterance[]) =>
RULES.map((r) => ({ id: r.id, hit: match(r.id, u) })).find((x) => x.hit) ?? null;Recording which rule fired is the part worth copying. Counting the distribution of fired rules per week turns a shift in how clinicians speak into a visible signal — you see the distribution move before the accuracy figure drops.
Values derived from a calendar belong in code
Settling the ranking is not enough on its own. Handing the model a resolved date and asking it to use that date still produces runs where a different date appears, because instructions inside a long prompt are sometimes skipped.
So code substitutes the date and relative expression tokens after the model has written the record.
Scoping the rewrite to tokens is the point. Replacing the whole line was on the table, but that variant can erase a clinician’s own annotation or a stated time. Broad machine rewriting of stored clinical text is the operation we most want to avoid.
Any such routine gets a safety valve, and how you count it matters more than the threshold you pick.
const out = replaceDateWords(line, decided); // only date and relative-expression tokens
// 🔴 Count the substitutions actually made, not the candidates found.
// Counting candidates makes the guard trip on the records that were
// already correct — the cleaner the input, the more likely it is discarded.
if (out.replacedCount > MAX_REPLACEMENTS) {
logger.warn('substitution count above limit; keeping the generated text', {
replaced: out.replacedCount, // 🔴 never the text itself
});
return line;
}Input validation follows the same principle in reverse: reject dangerous characters rather than allow-listing acceptable ones. An allow-list built from assumptions (“names are kanji and kana”) rejects real data — production names included compatibility variants of kanji and full-width symbols.
Nine rounds, all 56 recordings, scored by machine
Every design change was re-run against all 56 production recordings, every time. Rather than listening to each one, we push them through the same evaluation logic the product uses and score agreement with the human answer automatically.
| Round | Passing | What it surfaced |
|---|---|---|
| 1 | 20 of 36 | Confirmed that discarding stated dates was the main cause |
| 2 | 31 of 36 | A particle-separated phrasing was not being parsed at all |
| 4 | 35 of 36 | Unrelated appointments in the same consultation still leaked through |
| 6 | 54 of 56 | The last two vary because transcription itself varies per run |
Two findings mattered more than the scores. The first: the phrasing we were missing accounted for roughly 40% of production cases. Japanese speakers often insert a particle between month and day, and 59 of 161 production cases used that form exclusively. No amount of curating a test set gets you to that proportion — only production does.
The second: transcribing identical audio does not produce identical text. Two of the 56 recordings alternated between “next month” and “October 1st, Thursday” depending on the run. That is why we never conclude from a single scoring round.
The rounds did not stop there. After acting on code review, we re-ran the same 56 recordings three more times. Round nine: 53 passing, 2 informational, 1 failing on transcription drift, and zero disagreements with the human answer.
We also verify the guards themselves by deliberately breaking them. Seventeen mutations were introduced one at a time; sixteen produced a failing check.
The one that stayed green taught us the most. That check measured whether a call site appeared near a condition, by character distance. Move the call outside the condition it was supposed to protect and the check still passed. Review caught it; we rewrote the check to extract the block by bracket depth, then confirmed the same mutation turns it red. All seven mutations added after review failed correctly.
Restoring and confirming green is not optional. If it stays red, something other than your mutation is failing. Our injection tool asserts that the target string was actually present before editing, so “never broke it” can never be scored as “passed.”
A check existing and a check actually stopping something are different claims. So are “it stopped” and “it stopped for the reason you intended.”
After release, score on what clinicians corrected
This is the part we care most about. Scoring only before release means the side effects of your next design change stay invisible.
| Test recordings passed | Lines clinicians corrected | |
|---|---|---|
| What it measures | Accuracy against a curated set | What happened in real clinics |
| Failures it finds | Only the shapes already collected | Shapes nobody anticipated |
| Who authors truth | A person, one recording at a time | The correction itself is the truth |
| When it runs | Before release only | Every week, indefinitely |
Because clinical records keep a version history, adjacent versions can be compared to count how often only the next-appointment line changed. The shape of it:
-- Pair adjacent versions, count the ones where only that line changed
-- (column names simplified)
with pairs as (
select record_id,
next_appointment_line as now_line,
lag(next_appointment_line) over (
partition by record_id order by recorded_at
) as prev_line
from examination_record_history
where recorded_at >= now() - interval '7 days'
)
select count(*) filter (where prev_line is distinct from now_line) as edited,
count(*) as total -- 🔴 always return the denominator
from pairs
where prev_line is not null;A read-only database role runs it. We did not take that on trust: we issued a write over that connection and confirmed the database refused it with the expected error code. The guarantee that a measuring tool cannot damage records belongs in the connection, not in the tool’s own code.
The denominator ships with every number so a quiet week is never mistaken for an improvement. And because small denominators produce loud rates, we read the interval rather than the point estimate.
// Rate interval — wider the smaller the denominator
const band = (k: number, n: number, z = 1.96) => {
const p = k / n, d = 1 + (z * z) / n;
const c = p + (z * z) / (2 * n);
const m = z * Math.sqrt((p * (1 - p) + (z * z) / (4 * n)) / n);
return [(c - m) / d, (c + m) / d] as const;
};
// 9 of 77 → roughly 6%–21% 5 of 40 → roughly 5%–26%
// As point estimates, 11.7% and 12.5% look like a decline. The intervals overlap heavily.The mechanism paid for itself immediately. The first aggregation reported 76 of 77 records corrected — 98.7%, against a known measurement of 9.6%. An order of magnitude apart.
The cause was in the reader, not the data. The model’s original output was being wrapped under a second heading, so the routine that locates the line always answered “no line present,” and every record counted as corrected. After the fix: 9 of 77, or 11.7%. The scorer itself was broken.
Two things became permanent as a result. First, tests for the reading logic — inject the re-wrapping mutation and the tests go red; restore it and they go green. Second, a standing habit of comparing each weekly figure against a known measurement by order of magnitude. Weekly numbers arrive whether or not they are correct, so a broken scorer keeps reporting confidently. Measurement infrastructure needs the same scrutiny as the thing it measures.
On transcription varying between runs: we measured where that variance comes from separately, in When LLM Output Drifts, Suspect the Prompt Itself. The two together make the evaluation easier to design.
The practices, condensed
| Situation | Practice |
|---|---|
| Building an evaluation set | Derive it from production records a human corrected — both answers are already there |
| Writing decision rules | Order them by strength of evidence and record which rule fired; the distribution is an early warning |
| Rewriting generated text | Scope it to tokens. Count the substitutions made, never the candidates found |
| Validating input | Reject dangerous characters instead of allow-listing acceptable ones |
| Adding a guard | Break it, confirm red, restore, confirm green. The mutation that stays green is the weakness |
| Reporting a rate | Always publish the denominator, and read the interval |
| Building the aggregation | Run it on a read-only connection and verify the write is refused |
| Trusting a score | Compare it against a known measurement by order of magnitude. Scorers break too |
Summary
Three lines we hold when evaluating clinical AI accuracy:
- Let the clinic author ground truth. A corrected record already contains both the model’s answer and the human’s
- Keep derived values out of the model. Calendar arithmetic goes in code, and the rewrite stays scoped to tokens
- Keep measuring after release. Every design change should report back in next week’s number
Most clinics have a field somebody rewrites on nearly every record. That is countable, and once counted it can be designed away. We run exactly this loop on our own records each week, and the same shape can be built around an existing in-house template. If you would like to see how far SOAP documentation gets as a draft, showing us the format you use today is enough of a starting point — we are happy to talk.
Frequently asked questions
Why not simply collect more test recordings?
More of the same kind of recording will not surface failures outside that set. We score against production audio that includes cases a clinician already corrected. Such a record holds both the model output and the human answer, so we never author ground truth.
How is the human correction rate measured?
Clinical records keep a version history, so we compare adjacent versions and count how often only the next-appointment line changed. A read-only database role runs it, and the weekly summary always carries the denominator, so a quiet week is never read as an improvement.
Why not just instruct the model to use a given date?
We tried. Instructions buried in a long prompt are sometimes ignored, and the model still writes a different date on some runs. So date and relative expression tokens are substituted after generation, leaving times, notes and the clinical narrative untouched.
