Tech blog
Recovering a clinical note when the AI stops short
Contents
- The vocabulary, first
- The repetition was reuse, not similarity
- Input length was not the cause
- What did not work
- Why a model keeps writing when it has nothing to say
- Guarding one field moves the problem next door
- What worked was one rule that names no field
- Four layers, so the record survives a bad day
- Where it truncates changes what the fragment is worth
- The value you can parse but cannot store
- The safety net nearly destroyed the records
- Finding the recovered notes afterwards
- How we verified it
Sometimes the transcript succeeds and the clinical note never appears. The LLM starts writing, never stops, hits the output ceiling, and the record is left empty. For the nurse who recorded the visit, the outcome is simple: write it again by hand.
We reproduced this 658 times across 29 conditions in a staging environment. Two things came out of it. The guard we had already shipped was not stopping the problem, only moving it. And almost everything the internet recommends for repetition loops did nothing in our measurements.
This article lists every condition we tested, then explains how we decide which parts of a truncated note are safe to keep. The Japanese version is at 医療AIが記録を作りきれなかったとき、どこまで復旧できるか.
The vocabulary, first
| Term | What it means here |
|---|---|
| Token | The unit the model reads and writes. Japanese text is split roughly every one or two characters |
| Output ceiling | The maximum tokens one generation may produce. Reaching it ends the response mid-sentence |
| Truncation | The state of having hit that ceiling. The API reports it as a finish reason, not as an error |
| Structured output | Declaring the shape of the answer in advance, field by field, so it can be parsed reliably |
| Runaway repetition | The model rephrasing the same content indefinitely. It never errors; it simply fills the ceiling |
The repetition was reuse, not similarity
We counted the sentences inside a runaway field. There were 399 sentences and 10 distinct ones — a uniqueness rate of 2.5%.
| Scope | Sentences | Distinct | Uniqueness |
|---|---|---|---|
| The runaway field | 399 | 10 | 2.5% |
| The 24 fields written just before it, same record | — | — | 96.9–100% |
The same response contains healthy fields and one pathological field. The model does not degrade globally. It falls into the loop when it enters a particular field.
Here is how a truncated response actually ends (synthetic text, no patient data):
{
"schema_version": 1,
"observation": {
"general_condition": "Calm, settled expression throughout the visit.",
"sleep": "Bedtime is consistent and no complaint of waking during the night was raised. There were no specific complaints about sleep, and night-time rest appears to be adequate. No particular problems were described regarding sleep from bedtime to waking, so sleep is considered stable. Regarding difficulties with night-time sleepLook at the last field. The same statement, three times, reworded. The fourth rewording hits the ceiling mid-sentence, with no closing brace. At this point not one care or evaluation field has been written.
Input length was not the cause
| Transcript length | Output tokens on the runaway |
|---|---|
| 2,584 chars | 5,983 |
| 6,140 chars | 5,984 |
| 8,932 chars | 5,985 |
| 11,599 chars | 5,984 |
A fourfold difference in input produced the same output size. Healthy completions, by contrast, used a median of 1,291 tokens and a 95th percentile of 1,510 — the ceiling of 6,000 already sits at four times the p95. This is not a capacity problem. It is a stopping problem.
What did not work
Every row below was measured on the same audio, ten to thirty runs each.
| What we tried | Result | Why it failed |
|---|---|---|
| Raise the output ceiling | Rejected | Already 4× p95. Raising it lets the runaway finish inside the ceiling, so it returns as success and is saved without any signal |
| Declare per-field maxLength in the schema | 8 of 10 truncated | The schema declares shape, not a stopping rule |
| Frequency / presence penalties | 10 of 10 truncated | Some values were rejected by the API outright |
| Change the reasoning effort setting | 83% low, 67% high | Improves, does not solve |
| Move to a newer model generation | 5 of 10 truncated | And normal output became 2.7× more verbose |
| Allow the model to omit a field | 5% → 40% | Adding a decision made the deliberation longer, not shorter |
| Halve the number of fields | Still 83% | Field count was never the driver |
| Raise the temperature | No meaningful change | See below |
Temperature deserves its own table, because it is the first thing everyone reaches for.
| Output variability | Share of runs that ran away |
|---|---|
| Low (0.1) | Roughly unchanged |
| Medium (0.6) | Roughly unchanged |
| High (0.9) | Roughly unchanged |
Temperature barely moved this failure mode. Raising it on retry is still worthwhile — it avoids landing in the same probability well twice — but it is not the fix.
Why a model keeps writing when it has nothing to say
It helps to be concrete about the mechanism, because it explains why the successful fix looks so unimpressive.
Generation is autoregressive: each token is chosen from a distribution conditioned on everything written so far. Two properties of that loop matter here.
| Property | Consequence in a thin field |
|---|---|
The prompt says the field must be filled, and required in the schema means an empty string is the only alternative to prose |
Stopping is not a cheap option for the model |
| Whatever was just written is the strongest context for what comes next | A rephrasing is always the highest-probability continuation |
So the model enters a field with one sentence of material, is told to write in detail, cannot leave it empty, and the nearest plausible continuation is a restatement of what it just wrote. That restatement then becomes the context for the next one. The loop is not a bug in sampling; it is the most probable path given the instructions we wrote.
This is also why a maxLength in the schema does not rescue it. Constrained decoding shapes which tokens are legal, not which are likely, and the loop stays inside the grammar the whole time.
Guarding one field moves the problem next door
We had previously found this loop in one field (family involvement) and added an instruction naming that field and capping it. The field went quiet.
This time we measured again. The runaway had moved wholesale into the neighbouring sleep field.
| Condition | Runaways in the family field | Runaways in the sleep field |
|---|---|---|
| Before the field-specific guard | 59.7% of all runaways | 39.5% of all runaways |
| After guarding only the family field | 0 | 18 of 18 |
In healthy records these fields average 41 characters (sleep) and 106 characters (family). Both are fields where a single sentence in the conversation is normal. Any field where “write in detail” meets “almost nothing was said” behaves the same way. Naming fields in the guard is whack-a-mole.
What worked was one rule that names no field
| Condition | Runaways on the same audio |
|---|---|
| Before | 10 of 10 |
| After the global stopping rule | 0 of 100 (95% upper bound 3.0%) |
The rule has three parts, and we believe the third is doing most of the work:
- Every field must finish within a stated length and sentence count.
- A field with little material is correct to end short.
- Rephrasing to fill space, padding with generalities, and continuing to deliberate about whether to write are all forbidden.
Reading the runaway output makes the third one obvious: the text pouring into the record was the model reasoning about whether this content belonged in this field. With no material to write, it started deliberating instead of stopping.
Placement matters too. The same sentence placed at the start of the instructions left 8% of 25 runs truncated, against 0 of 100 when placed at the end. An instruction is not a constant; where you put it is part of the experiment.
Four layers, so the record survives a bad day
Even with the loop gone, a long visit can still reach the ceiling. Rather than one fix, the note is protected by four layers, each cheaper in information than the last.
Layer three is the interesting one, because the critical decision is what counts as finished.
// Accept a scalar only when the delimiter that must follow it
// was actually consumed (',' or '}'), never merely "input remains"
skipWs(c);
if (c.i >= c.s.length || (c.s[c.i] !== ',' && c.s[c.i] !== '}')) {
if (!v.isContainer && c.truncatedAt === null) c.truncatedAt = childPath;
return { complete: false, value: obj, isContainer: true };
}
obj[key] = v.value; // acceptance is confirmed only hereThe decision is this small:
We never inspect the text. An earlier guard in this system did judge by content — it detected repetition and blanked the field — and it deleted a patient’s genuine repeated complaint as a duplicate. In clinical records, the same words recurring can itself be the finding.
Our first implementation of this rule did not actually honour it. The comment said “only when the delimiter was consumed”; the code checked only that input remained, so "slept well"x was accepted. Review caught it. A stated invariant is not an implemented invariant.
A recovered record looks like this. Fields the model never reached are absent — not empty strings.
{
"schema_version": 1,
"observation": { "sleep": "…", "mental": { "mood": "…" }, "risk": { … } },
"care": {
"conversation_summary": "…",
"interventions": []
// family_engagement and education are simply not here
},
"evaluation": { "plan_until_next": [] }
// neither are assessment, doctor_report or next_visit_at
}Filling them with a placeholder would erase the difference between observed, nothing to report and never reached. Absence keeps both machine-detectable and visible on screen as not written.
We do restore containers — the arrays and nested objects the UI dereferences without checking — because a missing container crashes the screen while a missing leaf does not. We never restore a leaf. The next-visit date in particular is left absent: inventing it would mean recording “no follow-up planned”.
Where it truncates changes what the fragment is worth
Structured output is produced in schema order, so the truncation point is decided by the order of your fields.
| Truncated | Fields recovered | Saved? |
|---|---|---|
| After the conversation summary (59.7% of cases) | Most, including the summary | Yes |
| Before it (39.5% of cases) | 8 of 29 | No |
A note whose summary is missing is a note where almost every field is blank. A reader cannot tell “observed, nothing to report” from “the model never got here”. So we save only when the summary was recovered.
The value you can parse but cannot store
Truncated output is, by definition, broken JSON — and parsing it leniently produced a value that reads fine and cannot be stored.
A \uXXXX escape assumes four hex digits. Truncation cuts anywhere, so \uZZ shows up. Converting without validating the digits yields NaN, and String.fromCharCode(NaN) returns U+0000.
const hex = c.s.slice(c.i + 2, c.i + 6);
if (!/^[0-9a-fA-F]{4}$/.test(hex)) return { complete: false, value: out, isContainer: false };
const cp = parseInt(hex, 16);
if (cp === 0) return { complete: false, value: out, isContainer: false }; // jsonb cannot hold NUL
out += String.fromCharCode(cp);PostgreSQL’s jsonb rejects NUL outright — we confirmed 22P05 unsupported Unicode escape sequence against a real database. Without the check, the save would fail precisely on the runs where recovery succeeded: the exact inversion of the feature’s purpose. The same unvalidated advance also swallowed the closing quote, merging the next field into the current value and deleting that field.
We also capped nesting depth. A recursive-descent parser exhausts the stack around depth 5,000, and this code is the last line of defence — throwing here would lose the remaining retries as well. Past the cap we return “not readable” quietly.
The safety net nearly destroyed the records
Review found one more thing. Our “save what was written” step could overwrite a complete note.
When a user asks for re-analysis, the existing overwrite protection is deliberately bypassed — they asked for it. Add partial recovery underneath, and a finished note can be replaced by a fragment. We now compare before saving, by counting non-empty strings:
| Existing note | Recovered note | Decision |
|---|---|---|
| 0 (nothing written yet) | 31 | Save |
| 37 (complete) | 32 | Keep the existing note |
Consent to overwrite is not consent to be overwritten by less. When you add a safety net, count what is already in the place you are about to write.
Finding the recovered notes afterwards
We deliberately did not add a marker field to recovered records. A marker would have to live in the shared schema, which would have forced a redeploy of every function that imports it, and our own output guard would have stripped a path-like string from the record anyway.
Instead the absence is the signal. Two fields are present in every complete note and absent in a truncated one, so a two-field AND identifies recovered records exactly:
SELECT count(*) FROM nursing_records
WHERE NOT (record_data->'care' ? 'family_engagement')
AND NOT (record_data->'evaluation' ? 'assessment');We checked this against production before shipping: 0 of 372 existing records matched, so anything matching from now on is genuinely a recovered note. A single-field test would not have survived — under the new prompt one other field turned out to be absent in 0 of 7 healthy records where it had been absent in 77 of 288 under the old one. Absence is a property of the prompt, so pin it to more than one field.
How we verified it
| Method | What it covers | Result |
|---|---|---|
| Unit tests | Malformed escapes, cut digits, duplicate keys, deep nesting, code fences | 517 passing |
| Mutation testing | Break each guard one at a time, confirm the suite goes red | 11 of 11 went red |
| Live re-measurement | Same audio through the real pipeline after the fix | 0 truncations |
The second matters most. A passing suite and a meaningful suite are different things: if breaking the guard leaves the tests green, the tests were never watching it. One of ours stayed green the first time, and we rewrote it.
MENTRA builds clinical documentation for psychiatric care in Japan, where the note has to survive the bad days as well as the good ones. The Japanese version of this article is at 医療AIが記録を作りきれなかったとき、どこまで復旧できるか, and our visiting-nurse documentation page describes where this runs.
Frequently asked questions
Is the note truncated because the recording was too long?
Usually not. Across our runs the transcript ranged from 2,584 to 11,599 characters, and every runaway pinned at 5,983 to 5,985 output tokens. Output length was not a function of input length. Suspect an instruction telling the model to write in detail into a field with nothing to say.
Can a maxLength in the response schema prevent this?
It did not in our measurements. With a per-field character limit declared in the structured output schema, eight of ten runs still hit the token ceiling. Treat the schema as a declaration of shape, not as a mechanism that stops generation.
Is it safe to save a partially written clinical note?
Only if you decide what counts as finished by grammar. We accept a field solely when the parser consumed the closing delimiter and the following structural token, and never inspect the text. Judging by content is how an earlier guard deleted a real repeated complaint as a duplicate.
Can staff tell a recovered note from a complete one?
Yes. Fields the model never reached are absent rather than filled with a placeholder, so the screen shows them as not written and they stay machine-detectable. A default string would erase the difference between not observed and not reached.
