Tech blog

Recovering a clinical note when the AI stops short

Contents
  1. The vocabulary, first
  2. The repetition was reuse, not similarity
  3. Input length was not the cause
  4. What did not work
  5. Why a model keeps writing when it has nothing to say
  6. Guarding one field moves the problem next door
  7. What worked was one rule that names no field
  8. Four layers, so the record survives a bad day
  9. Where it truncates changes what the fragment is worth
  10. The value you can parse but cannot store
  11. The safety net nearly destroyed the records
  12. Finding the recovered notes afterwards
  13. How we verified it

Sometimes the transcript succeeds and the clinical note never appears. The LLM starts writing, never stops, hits the output ceiling, and the record is left empty. For the nurse who recorded the visit, the outcome is simple: write it again by hand.

We reproduced this 658 times across 29 conditions in a staging environment. Two things came out of it. The guard we had already shipped was not stopping the problem, only moving it. And almost everything the internet recommends for repetition loops did nothing in our measurements.

This article lists every condition we tested, then explains how we decide which parts of a truncated note are safe to keep. The Japanese version is at 医療AIが記録を作りきれなかったとき、どこまで復旧できるか.

The vocabulary, first

Term What it means here
Token The unit the model reads and writes. Japanese text is split roughly every one or two characters
Output ceiling The maximum tokens one generation may produce. Reaching it ends the response mid-sentence
Truncation The state of having hit that ceiling. The API reports it as a finish reason, not as an error
Structured output Declaring the shape of the answer in advance, field by field, so it can be parsed reliably
Runaway repetition The model rephrasing the same content indefinitely. It never errors; it simply fills the ceiling

The repetition was reuse, not similarity

We counted the sentences inside a runaway field. There were 399 sentences and 10 distinct ones — a uniqueness rate of 2.5%.

Scope Sentences Distinct Uniqueness
The runaway field 399 10 2.5%
The 24 fields written just before it, same record — — 96.9–100%

The same response contains healthy fields and one pathological field. The model does not degrade globally. It falls into the loop when it enters a particular field.

Here is how a truncated response actually ends (synthetic text, no patient data):

{
  "schema_version": 1,
  "observation": {
    "general_condition": "Calm, settled expression throughout the visit.",
    "sleep": "Bedtime is consistent and no complaint of waking during the night was raised. There were no specific complaints about sleep, and night-time rest appears to be adequate. No particular problems were described regarding sleep from bedtime to waking, so sleep is considered stable. Regarding difficulties with night-time sleep

Look at the last field. The same statement, three times, reworded. The fourth rewording hits the ceiling mid-sentence, with no closing brace. At this point not one care or evaluation field has been written.

Input length was not the cause

Transcript length Output tokens on the runaway
2,584 chars 5,983
6,140 chars 5,984
8,932 chars 5,985
11,599 chars 5,984

A fourfold difference in input produced the same output size. Healthy completions, by contrast, used a median of 1,291 tokens and a 95th percentile of 1,510 — the ceiling of 6,000 already sits at four times the p95. This is not a capacity problem. It is a stopping problem.

What did not work

Every row below was measured on the same audio, ten to thirty runs each.

What we tried Result Why it failed
Raise the output ceiling Rejected Already 4× p95. Raising it lets the runaway finish inside the ceiling, so it returns as success and is saved without any signal
Declare per-field maxLength in the schema 8 of 10 truncated The schema declares shape, not a stopping rule
Frequency / presence penalties 10 of 10 truncated Some values were rejected by the API outright
Change the reasoning effort setting 83% low, 67% high Improves, does not solve
Move to a newer model generation 5 of 10 truncated And normal output became 2.7× more verbose
Allow the model to omit a field 5% → 40% Adding a decision made the deliberation longer, not shorter
Halve the number of fields Still 83% Field count was never the driver
Raise the temperature No meaningful change See below

Temperature deserves its own table, because it is the first thing everyone reaches for.

Output variability Share of runs that ran away
Low (0.1) Roughly unchanged
Medium (0.6) Roughly unchanged
High (0.9) Roughly unchanged

Temperature barely moved this failure mode. Raising it on retry is still worthwhile — it avoids landing in the same probability well twice — but it is not the fix.

Why a model keeps writing when it has nothing to say

It helps to be concrete about the mechanism, because it explains why the successful fix looks so unimpressive.

Generation is autoregressive: each token is chosen from a distribution conditioned on everything written so far. Two properties of that loop matter here.

Property Consequence in a thin field
The prompt says the field must be filled, and required in the schema means an empty string is the only alternative to prose Stopping is not a cheap option for the model
Whatever was just written is the strongest context for what comes next A rephrasing is always the highest-probability continuation

So the model enters a field with one sentence of material, is told to write in detail, cannot leave it empty, and the nearest plausible continuation is a restatement of what it just wrote. That restatement then becomes the context for the next one. The loop is not a bug in sampling; it is the most probable path given the instructions we wrote.

This is also why a maxLength in the schema does not rescue it. Constrained decoding shapes which tokens are legal, not which are likely, and the loop stays inside the grammar the whole time.

Guarding one field moves the problem next door

We had previously found this loop in one field (family involvement) and added an instruction naming that field and capping it. The field went quiet.

This time we measured again. The runaway had moved wholesale into the neighbouring sleep field.

Family field thin material, runs away Guard that field named, capped Family field 0 runaways Sleep field now does it thin material is not unique to one field
Name a field in the guard and the runaway moves to the next field with the same shape
Condition Runaways in the family field Runaways in the sleep field
Before the field-specific guard 59.7% of all runaways 39.5% of all runaways
After guarding only the family field 0 18 of 18

In healthy records these fields average 41 characters (sleep) and 106 characters (family). Both are fields where a single sentence in the conversation is normal. Any field where “write in detail” meets “almost nothing was said” behaves the same way. Naming fields in the guard is whack-a-mole.

What worked was one rule that names no field

Condition Runaways on the same audio
Before 10 of 10
After the global stopping rule 0 of 100 (95% upper bound 3.0%)

The rule has three parts, and we believe the third is doing most of the work:

  1. Every field must finish within a stated length and sentence count.
  2. A field with little material is correct to end short.
  3. Rephrasing to fill space, padding with generalities, and continuing to deliberate about whether to write are all forbidden.

Reading the runaway output makes the third one obvious: the text pouring into the record was the model reasoning about whether this content belonged in this field. With no material to write, it started deliberating instead of stopping.

Placement matters too. The same sentence placed at the start of the instructions left 8% of 25 runs truncated, against 0 of 100 when placed at the end. An instruction is not a constant; where you put it is part of the experiment.

Four layers, so the record survives a bad day

Even with the loop gone, a long visit can still reach the ceiling. Rather than one fix, the note is protected by four layers, each cheaper in information than the last.

1. Do not run away global stopping rule 2. Retry raise temperature 3. Keep what exists grammar decides 4. Protect never overwrite stopping at an earlier layer is always better
Each layer is a fallback for the one before it, and keeps less information

Layer three is the interesting one, because the critical decision is what counts as finished.

// Accept a scalar only when the delimiter that must follow it
// was actually consumed (',' or '}'), never merely "input remains"
skipWs(c);
if (c.i >= c.s.length || (c.s[c.i] !== ',' && c.s[c.i] !== '}')) {
  if (!v.isContainer && c.truncatedAt === null) c.truncatedAt = childPath;
  return { complete: false, value: obj, isContainer: true };
}
obj[key] = v.value;            // acceptance is confirmed only here

The decision is this small:

Value parsed Next char , or } ? actually consumed Accept completion proven no Discard this is the cut point
Acceptance is decided by form alone. The text is never read

We never inspect the text. An earlier guard in this system did judge by content — it detected repetition and blanked the field — and it deleted a patient’s genuine repeated complaint as a duplicate. In clinical records, the same words recurring can itself be the finding.

Our first implementation of this rule did not actually honour it. The comment said “only when the delimiter was consumed”; the code checked only that input remained, so "slept well"x was accepted. Review caught it. A stated invariant is not an implemented invariant.

A recovered record looks like this. Fields the model never reached are absent — not empty strings.

{
  "schema_version": 1,
  "observation": { "sleep": "…", "mental": { "mood": "…" }, "risk": { … } },
  "care": {
    "conversation_summary": "…",
    "interventions": []
    // family_engagement and education are simply not here
  },
  "evaluation": { "plan_until_next": [] }
    // neither are assessment, doctor_report or next_visit_at
}

Filling them with a placeholder would erase the difference between observed, nothing to report and never reached. Absence keeps both machine-detectable and visible on screen as not written.

We do restore containers — the arrays and nested objects the UI dereferences without checking — because a missing container crashes the screen while a missing leaf does not. We never restore a leaf. The next-visit date in particular is left absent: inventing it would mean recording “no follow-up planned”.

Where it truncates changes what the fragment is worth

Structured output is produced in schema order, so the truncation point is decided by the order of your fields.

Observation fields sleep lives here Conversation summary the spine of the note Care and evaluation fields family involvement, plan truncate here: 8 of 29 fields, do not save truncate here: most fields survive, save
The truncation point is a property of your schema order, not of the model
Truncated Fields recovered Saved?
After the conversation summary (59.7% of cases) Most, including the summary Yes
Before it (39.5% of cases) 8 of 29 No

A note whose summary is missing is a note where almost every field is blank. A reader cannot tell “observed, nothing to report” from “the model never got here”. So we save only when the summary was recovered.

The value you can parse but cannot store

Truncated output is, by definition, broken JSON — and parsing it leniently produced a value that reads fine and cannot be stored.

A \uXXXX escape assumes four hex digits. Truncation cuts anywhere, so \uZZ shows up. Converting without validating the digits yields NaN, and String.fromCharCode(NaN) returns U+0000.

const hex = c.s.slice(c.i + 2, c.i + 6);
if (!/^[0-9a-fA-F]{4}$/.test(hex)) return { complete: false, value: out, isContainer: false };
const cp = parseInt(hex, 16);
if (cp === 0) return { complete: false, value: out, isContainer: false };  // jsonb cannot hold NUL
out += String.fromCharCode(cp);

PostgreSQL’s jsonb rejects NUL outright — we confirmed 22P05 unsupported Unicode escape sequence against a real database. Without the check, the save would fail precisely on the runs where recovery succeeded: the exact inversion of the feature’s purpose. The same unvalidated advance also swallowed the closing quote, merging the next field into the current value and deleting that field.

We also capped nesting depth. A recursive-descent parser exhausts the stack around depth 5,000, and this code is the last line of defence — throwing here would lose the remaining retries as well. Past the cap we return “not readable” quietly.

The safety net nearly destroyed the records

Review found one more thing. Our “save what was written” step could overwrite a complete note.

When a user asks for re-analysis, the existing overwrite protection is deliberately bypassed — they asked for it. Add partial recovery underneath, and a finished note can be replaced by a fragment. We now compare before saving, by counting non-empty strings:

Existing note Recovered note Decision
0 (nothing written yet) 31 Save
37 (complete) 32 Keep the existing note

Consent to overwrite is not consent to be overwritten by less. When you add a safety net, count what is already in the place you are about to write.

Finding the recovered notes afterwards

We deliberately did not add a marker field to recovered records. A marker would have to live in the shared schema, which would have forced a redeploy of every function that imports it, and our own output guard would have stripped a path-like string from the record anyway.

Instead the absence is the signal. Two fields are present in every complete note and absent in a truncated one, so a two-field AND identifies recovered records exactly:

SELECT count(*) FROM nursing_records
WHERE NOT (record_data->'care'       ? 'family_engagement')
  AND NOT (record_data->'evaluation' ? 'assessment');

We checked this against production before shipping: 0 of 372 existing records matched, so anything matching from now on is genuinely a recovered note. A single-field test would not have survived — under the new prompt one other field turned out to be absent in 0 of 7 healthy records where it had been absent in 77 of 288 under the old one. Absence is a property of the prompt, so pin it to more than one field.

How we verified it

Method What it covers Result
Unit tests Malformed escapes, cut digits, duplicate keys, deep nesting, code fences 517 passing
Mutation testing Break each guard one at a time, confirm the suite goes red 11 of 11 went red
Live re-measurement Same audio through the real pipeline after the fix 0 truncations

The second matters most. A passing suite and a meaningful suite are different things: if breaking the guard leaves the tests green, the tests were never watching it. One of ours stayed green the first time, and we rewrote it.

MENTRA builds clinical documentation for psychiatric care in Japan, where the note has to survive the bad days as well as the good ones. The Japanese version of this article is at 医療AIが記録を作りきれなかったとき、どこまで復旧できるか, and our visiting-nurse documentation page describes where this runs.

Frequently asked questions

Is the note truncated because the recording was too long?

Usually not. Across our runs the transcript ranged from 2,584 to 11,599 characters, and every runaway pinned at 5,983 to 5,985 output tokens. Output length was not a function of input length. Suspect an instruction telling the model to write in detail into a field with nothing to say.

Can a maxLength in the response schema prevent this?

It did not in our measurements. With a per-field character limit declared in the structured output schema, eight of ten runs still hit the token ceiling. Treat the schema as a declaration of shape, not as a mechanism that stops generation.

Is it safe to save a partially written clinical note?

Only if you decide what counts as finished by grammar. We accept a field solely when the parser consumed the closing delimiter and the following structural token, and never inspect the text. Judging by content is how an earlier guard deleted a real repeated complaint as a duplicate.

Can staff tell a recovered note from a complete one?

Yes. Fields the model never reached are absent rather than filled with a placeholder, so the screen shows them as not written and they stay machine-detectable. A default string would erase the difference between not observed and not reached.

↑ Back to top

まずは、30分だけ話を聞かせてください。

今の記録の流れを伺うところからで大丈夫です。資料の準備は不要です。

精神科病院からメンタルクリニックまで、規模を問わず。院内の運用に合わせた調整もご相談ください。