Tech blog

Silence makes medical AI invent clinical conversations

Contents
  1. What silence actually returns
  2. Length tells you nothing
  3. This is not one vendor’s problem
  4. Asking the model to stay silent does not work
  5. So we measure the audio instead
  6. Never let “not measured” read as “fine”
  7. Put the guard where it cannot be skipped
  8. Measuring long recordings without decoding them
  9. What we would tell another team

This article is also available in Japanese.

When you build clinical documentation on top of a speech model, the frightening failure is not the model being wrong. It is the model being wrong in a way nobody can see.

Give it a recording that contains no sound at all. Intuitively it should say it heard nothing. It does not. It returns a fluent, well-formed clinical conversation that a reader cannot distinguish from a real one.

What silence actually returns

We measured it. We took audio containing no signal whatsoever and ran it through our production documentation pipeline. Repeating identical input returns cached results, so we varied only a name hint to make each run independent.

Model Runs Runs that invented a conversation Characters generated
Current production model 8 8 of 8 410 / 421 / 422 / 434 / 447 / 496 / 603 / 775
Previous production model 8 6 of 6 that completed 63 / 170 / 171 / 240 / 263 / 414

Every run that completed invented a clinical conversation. Here is one, translated from the Japanese, with a dummy patient name. The drug name is real and appeared verbatim in the output.

Mr. Yamada, hello. How have you been feeling lately? — Hello. Lately I keep waking up several times in the night and can’t sleep soundly. — So you’re having middle-of-the-night awakening. What about daytime sleepiness or fatigue? — … You’re currently taking Dayvigo. Do you feel it isn’t working as well?

Note what is present: a presenting complaint, the clinician’s follow-up question, a real prescription drug, and a check on its effect. It reads as an ordinary insomnia follow-up. Nothing in the text suggests it came from silence.

Silent recording no signal at all Documentation no error raised A complete clinical note symptoms, questions, real drug names saved like any other record
No stage fails, so the caller sees success.

No stage in that pipeline raises an error. Transcription succeeds. Every field of the template is populated. The record saves. The only signal a human has is the memory of never having had that conversation.

Length tells you nothing

The obvious first defence is to distrust short output. It does not work.

Look at the table again. From the same duration of the same silence, one run produced 63 characters and another produced 775. A twelvefold spread. The amount written is essentially unrelated to the input; it reflects how much the model felt like writing on that pass.

So no threshold exists, on character count or on characters per second. Any threshold you pick will flag a terse consultation with a quiet patient and wave through a verbose fabrication.

We tried other discriminators as well.

Candidate signal What it looks at Result
Output length or density How much the model wrote ❌ Same silence yields 63 or 775 characters. Unrelated to input
Bytes per second of the file Encoded data rate ❌ We hold silent recordings at 3,627 B/s and real speech at 1,072 B/s. The ordering inverts
Peak amplitude alone Loudest sample in the file ❌ A single click can push a silent recording to nearly full scale
Share of time below a level Portion under −35 dBFS ❌ Real but soft speech also measures as 100% silent
Switching models Use a different provider ❌ Both the current and previous models fabricated on every completed run
Failing thin output wholesale Error out when output is short ❌ On real recordings this caught a meaningful number of legitimate short notes

This is not one vendor’s problem

The phenomenon is documented repeatedly across speech AI. Every figure below is taken from the primary source.

Source What was measured Result
Barański et al. 2025, AGH Kraków 301,317 non-speech clips through Whisper 40.3% produced text
Same paper Real speech with silence appended 17.1% hallucinated. Padding alone causes it
Koenecke et al. 2024, ACM FAccT 13,140 speech segments from AphasiaBank (speakers with aphasia and a control group) 187 segments reliably hallucinated; 38% of hallucinations were harmful: violence, made-up names and health conditions, fake website links
Calm-Whisper 2025, Interspeech 8,732 environmental sound clips Whisper emitted text 99.97% of the time. It returned empty three times
When Silence Matters 2025, NTU Five seconds of unrelated silence appended for audio language models Accuracy fell consistently. Silence is not a neutral input

Clinical harm has followed. In May 2026 the Auditor General of Ontario procurement-tested 20 approved AI scribe vendors on simulated consultations and found inaccuracies in all twenty. Nine invented treatment plans that were never discussed. Twelve recorded medications different from what the physician prescribed.

In August 2026 ABC reported that an Australian patient’s post-operative letter from her specialist to her GP had acquired a drug history she had never had; she had noticed it in March. The part that matters most is not the fabrication. It is that the clinician could not determine where it came from. Without the original audio there is nothing left to check against.

So we do not delete the audio. When a record looks wrong, being able to listen again is the last line of defence against this class of error.

Asking the model to stay silent does not work

The next instinct is to instruct the model: if you hear nothing, return nothing. This has been tested and disproven, from three independent directions.

  1. With Whisper’s initial prompt, telling it not to invent words causes the instruction itself to appear in the transcript. The prompt enters the decoder as context, so you have simply given it one more thing to repeat.
  2. Measured on audio language models (NTU 2025), mitigation prompts improved results inconsistently and made some conditions worse. The paper concludes that a single explicit instruction is insufficient to counteract cross-modal interference systematically.
  3. Work on audio hallucination attacks found that chain-of-thought prompting helped against explicit attacks but not implicit ones, and in one setting raised the attack success rate from 68.74% to 82.90%. What helped substantially was training-time alignment (DPO).

Do not delegate hallucination safety to a prompt. Prompts get edited. Models get swapped. A defence that lives in the instruction text is not a defence.

So we measure the audio instead

We stopped trying to recognise fabrication and started making sure the model never sees silence. Front-end voice activity detection is also the strongest single mitigation in the literature: Whisper’s hallucination rate drops from 21.3% to 0.2% with Silero VAD in front of it.

Our rule is deliberately narrow.

Green peaks are intervals above the level. Gaps between them may be any length Measure every 1 ms louder of the two channels Sum the intervals above total, not longest run Under 1 s total: silent never reaches the model
Sum the voiced time. Requiring one continuous second rejects real conversation.

Total, not longest run. Our first implementation required one continuous second above the threshold. Run a genuine but heavily attenuated conversation through that and each word falls short of a second, so real speech is judged silent. Conversation is intermittent by nature; requiring continuity penalises the people who speak quietly.

The detector itself is a short function. It takes the RMS of one millisecond at a time, adds up the time spent above the threshold, and stops reading the moment the total reaches one second.

export const SILENCE_LEVEL_DBFS = -50;     // floor for "voiced"
export const SILENCE_MIN_VOICED_SEC = 1.0; // total time above it

const threshold = Math.pow(10, SILENCE_LEVEL_DBFS / 20);  // dBFS to amplitude
const win  = Math.round(sampleRate * SILENCE_RMS_WINDOW_MS / 1000);
const need = Math.round(sampleRate * SILENCE_MIN_VOICED_SEC);

for (let i = 0; i + win <= pcm.length; i += win) {
  let sum = 0;
  for (let j = i; j < i + win; j++) { const v = pcm[j]; sum += v * v; }
  if (Math.sqrt(sum / win) > threshold) {        // this window is voiced
    voicedSamples += win;
    if (voicedSamples >= need) return true;      // one second total, stop here
  }
}

The same function runs in the browser during recording and on the server after upload — fed from an AudioWorklet in one case and from the decoder in the other. A check in CI pins both the constants and the hash of that core block so the two cannot drift apart.

The two constants, −50 dBFS and one second in total, come from measurement rather than intuition. We tested four questions separately.

Question What we measured
Is −50 dBFS low enough? Silent files never cross it (peak −68.7 dBFS at most; isolated clicks cross but never last a second)
Does it catch quiet speech? The real recording with the least voiced time still had 11.6 seconds. There is a gap, not a gradient
Is one second the right length? It is the shortest length that still classifies a single click (peak −0.089 dBFS) as silence
Could we be stricter? Yes, and there is no reason to. The quietest real speech peaks at −15.2 dBFS, leaving 35 dB of headroom

Against our sample set: 20 of 20 silent recordings caught, 0 of 85 real recordings flagged. The guard is meant to catch recordings with no signal, and it must not go looking for quiet ones.

Never let “not measured” read as “fine”

The verdict has three values, not two: voiced, silent, and unmeasured.

type SilenceVerdict = 'voiced' | 'silent' | 'unmeasured';

type AudioCheck =
  | { verdict: 'voiced' }                                    // goes to the model
  | { verdict: 'silent';     reason: 'no_voiced_run' }       // stopped
  | { verdict: 'unmeasured'; reason: 'unsupported_format'    // never actually measured
                                   | 'no_audio_track'
                                   | 'decode_failed'
                                   | 'budget_exhausted' };

Storing the reason next to the verdict is what makes “how many unsupported containers did we see this month” a query rather than an investigation.

Unsupported container, no audio track, a decode failure, the scan budget running out — collapse these into “silent” and you block valid recordings; collapse them into “voiced” and you have passed something you never checked. Both are wrong, so we added the third value and recorded the reason alongside it. Afterwards it is possible to tell an unsupported format from an exhausted budget.

When two different causes produce the same outcome, nobody can ever trace back to the cause.

Put the guard where it cannot be skipped

One more design point mattered as much as the rule itself.

Five separate code paths fetched audio, one per record type. Writing a measureAudio() helper and calling it from all five means someone eventually forgets one call — and nothing reveals it until that record type produces a fabricated note.

When the same shape exists five times, fixing one leaves four copies as templates for the next change. What multiplies by copying gets fixed by copying too, and one copy is always missed. We have watched layered safeguards all miss the same single path when we measured how far our own checks actually reached.

So we collapsed the five fetches into a single function. Audio cannot be read without passing through it.

Consultation notes Nursing records Interview records Welfare visit records Conference minutes One audio fetch that measures sound Voiced: continue unmeasured continues too Silent: stop audio kept, record empty
A checkpoint you cannot route around is structure. One you are supposed to call is a convention.

When the verdict is silent we skip the model entirely, leave the record empty, and keep the audio. Staff are told the recording did not capture sound, and they can listen for themselves. We would rather have an absent record than an invented one.

Measuring long recordings without decoding them

Fully decoding a forty-minute recording just to answer one question is too slow, so we sample across the timeline. Two things went wrong here, and both are worth passing on.

First, where you sample matters more than how much. Our initial design decoded one continuous second at a fixed stride. Tested against a deliberately hard case — a 62-minute recording containing just 15 seconds of speech — it returned “silent” on three runs out of five. With a single window, if that one second falls in a gap, you miss.

Splitting the same budget into five windows of 0.2 seconds fixed it. Same total audio decoded, five out of five correct.

The whole policy reduces to deciding which frames to decode. The stride is chosen at runtime from the estimated duration and the remaining budget:

const SAMPLE_WINDOW_FRAMES = 50;   // 20ms x 50 = one second of audio per cycle
const SAMPLE_SUB_WINDOWS   = 5;    // placed as five separate 0.2s windows

export function shouldDecodeFrame(seen: number, stride: number): boolean {
  if (stride <= 1) return true;                  // no need to skip, read it all
  const sub = Math.round(SAMPLE_WINDOW_FRAMES / SAMPLE_SUB_WINDOWS);  // 0.2s = 10 frames
  return (seen % (stride * sub)) < sub;          // read 0.2s, skip the gap
}

False positives did not increase: a 62-minute pure silence and a 1h42m entirely silent recording both stayed silent, and across 89 real recordings not one verdict changed. Worst-case runtime was 1,343ms against a 1,500ms budget.

Sampling strategy Audio decoded 62 min recording, 15 s of speech Pure silence (negative control)
One 1-second window same 2 of 5 runs correct correctly silent
Five 0.2-second windows same 5 of 5 runs correct correctly silent

Second, the guard we added created a new way to lose data. We capped how much audio the scanner would read. In our runtime, however, a blob’s stream yields the entire blob as a single chunk, so the cap tripped on chunk one and any recording above the cap became unreadable in its entirety. The larger the recording, the more certain the loss.

We caught it before release, rewrote the reader to slice fixed-size ranges, and added tests using files just under and just over the limit. Then we deliberately broke each guard to confirm the tests actually turn red.

Every guard you add should come with one condition under which it fails. If you cannot name that condition, you have not finished designing the guard.

What we would tell another team

Two questions decide whether this class of error is survivable in your system. Can someone listen to the original audio afterwards? And when a record ends up empty, does anyone find out? If both answers are yes, every mistake remains traceable.

MENTRA is a medical AI built for psychiatry and psychosomatic medicine in Japan. All five of our record types — consultation notes, nursing records, interview records, welfare visit records and conference minutes — read audio through a single entry point that runs this check every time, and a scheduled job reports the results daily.

The full Japanese version of this article, with the same measurements, is here. If you are working on clinical documentation and want to compare notes on any of this, we would be glad to hear from you — get in touch.

Frequently asked questions

Does the guard also block quiet consultations or soft-spoken patients?

No. It only catches recordings with no signal at all. Across our real speech samples, the quietest one still had more than ten times the voiced time our threshold requires. Distant microphones, soft voices and long pauses all pass.

How much latency does measuring the audio add?

Tens of milliseconds in practice. Long recordings are never fully decoded. We sample across the timeline and stop reading the moment we find enough voiced audio, so a normal recording is decided within the first fraction of it.

What happens when the audio cannot be measured at all?

It becomes a third state, not a pass or a fail. Unsupported container, no audio track, decode failure and budget exhaustion are each recorded with their reason, and the audio still goes to the model. Collapsing unmeasured into either answer hides whether the check ever ran.

Why not just detect the fabrication afterwards?

Because the fabrication is fluent, complete and internally consistent. It fills every field of the template and passes every format check. Detection has to happen where the evidence still exists, which is the audio, not the text.

↑ Back to top

まずは、30分だけ話を聞かせてください。

今の記録の流れを伺うところからで大丈夫です。資料の準備は不要です。

精神科病院からメンタルクリニックまで、規模を問わず。院内の運用に合わせた調整もご相談ください。