Tech blog
When LLM Output Drifts, Suspect the Prompt Itself
Contents
- The vocabulary, first
- The measurement
- Temperature 0 is a sampling rule, not a determinism guarantee
- How to check this in practice
- What the 16 characters were
- Prompt ordering is a design decision
- The fix, and the trade-off we accepted
- Was the instruction ignored, or never delivered?
- The tooling behind the numbers
- What the rig looks like in production
- What to hold fixed when comparing
- The practices, condensed
- Summary
When teams put a large language model into a production workflow, the first knob they reach for is temperature. Setting it to 0 is widely assumed to mean “the same input gives the same output.” It does not.
We measured this on a real workload: turning recordings of a weekly clinical meeting into structured minutes. Changing 16 characters inside the prompt — characters that carry no meaning at all — made the same audio produce a different classification on every run.
This post covers what we measured, why it happens, and what has to stay fixed. A Japanese version is available at 生成AIの答えがぶれるとき、プロンプトのどこを疑うか.
The vocabulary, first
Everything below uses these terms. This table is enough to follow the measurements.
| Term | What it means |
|---|---|
| Token | The smallest unit the model reads and writes — finer than a word. Text is converted into a sequence of these before the model sees it |
| Logit | The raw score for how plausible each candidate is as the next token, before it becomes a probability |
| Softmax | The step that turns a row of logits into probabilities that sum to 1 |
| Temperature | How flat or peaked that probability distribution is. At 0 the peak wins outright (greedy decoding) |
| Greedy decoding | Always taking the highest-probability candidate. No draw is made, so temperature 0 is equivalent to this |
| Seed | Fixes the random draw used in sampling. With no draw happening at temperature 0, its reach is limited |
| Prefix caching | Reusing computed state across requests that share a leading prefix. A changed prefix means no reuse |
| Batch | The group of concurrent requests a serving stack computes together. Its composition changes with load |
| Thinking budget | How much internal deliberation the model spends before answering. More costs time and money, and pays off little on short, templated work |
| Structured output | Constraining the response to a declared schema of fields and types |
The measurement
The task: a weekly meeting at a psychiatric hospital produces minutes in a fixed format. Every report has to be filed into one of eight sub-sections. If the filing changes between runs, the document cannot be used as a formal record.
| Parameter | Value |
|---|---|
| Audio | 3 real meeting recordings + 3 synthesised ones with the same structure |
| Runs | 3 per recording, re-running analysis only against a fixed transcript |
| Checks | 26 per run (numeric agreement, section placement, formatting) + per-recording stability |
| Total | 468 checks |
| Temperature | 0 |
Holding the transcript fixed matters. Speech recognition varies slightly on every pass, so if you re-transcribe each time you cannot separate analysis variance from recognition variance. We froze the text and varied one thing.
“Stability” here means: run the same recording three times and check whether all eight sub-sections received the same content every time. One mismatch marks that recording unstable.
Scoring was automated from the start, and it carries one guard that turned out to matter more than the scoring itself: if the number of checks ever falls below 468, the scoring run fails.
We added it the hard way. Partway through, the runner silently dropped outputs and the denominator fell from 468 to 417. The pass rate rose purely because the denominator shrank, and the run read as an improvement. Pinning the denominator and failing loudly ended that class of misreading.
Temperature 0 is a sampling rule, not a determinism guarantee
Producing one token looks roughly like this:
- Convert the input text into a sequence of tokens
- Compute a logit for every candidate next token
- Softmax those logits into probabilities
- At temperature 0, take the highest one
The drift enters at step 2. Inference runs thousands of parallel reductions on a GPU. Floating-point addition is not associative:
(a + b) + c ≠ a + (b + c) // in general, for IEEE-754 floatsUsually that is invisible. It stops being invisible when the top two candidates are close:
| Candidate | Logit | Probability |
|---|---|---|
| File under “upcoming admissions” | 8.4213417 | 50.0002% |
| File under “items for discussion” | 8.4213409 | 49.9998% |
The gap is in the seventh decimal place. Change how the work is split across threads and that digit flips — and temperature 0 simply takes whichever is higher, so the section changes the instant the ranking does. In autoregressive generation, one different token changes the distribution for every token after it, so a difference of one word compounds into a different paragraph.
Worse, the split depends on batch composition. Your request can be identical while the server’s concurrent load is not, so the reduction order differs. That condition is invisible from the client side. What temperature 0 actually buys you is a narrow variance band, not repeatability.
How to check this in practice
- Look at the margin between the top two candidates. Most APIs can return per-token log probabilities. If the margin is small at the position where the classification is decided, that decision will keep drifting. Inspect the deciding token, not the whole response
- Run the same input at least three times and keep the agreement rate. A single run is a state, not a performance figure. We use “all three identical” as the pass condition
- Re-measure when the model version changes. When the provider’s stack rolls over, the variance profile changes with it. Record the version and never mix measurements across it
What the 16 characters were
The meeting format is editable by hospital staff through the UI. Concatenating that text straight into a prompt would let anyone write “ignore the instructions above” into a template field. So we wrapped the template data in a fence:
【FORMAT BEGIN #k3n8vq】
1. Reports
(1) Current census …
【FORMAT END #k3n8vq】#k3n8vq is the marker. With a constant marker, someone editing the template could simply type the closing line themselves and appear to escape the fence. So we changed it to a fresh random value on every request. As a security change, this is the right direction.
Then we re-measured.
| Fence strategy | Stable recordings | Runs with no leftover placeholders |
|---|---|---|
| Derived from content (before) | 6 / 6 | 36 / 36 |
| Random per request (after) | 1 / 3 | 8 / 9 |
| Derived per record (compromise) | 2 / 3 | 9 / 9 |
From 6/6 to 1/3. Same template, same transcript, same model, same temperature. The only difference was 16 characters of noise.
The mechanism is not subtle. A changing marker means a changing token sequence. We thought we were sending the same input three times; we were sending three different inputs, and the “close candidates flip” behaviour above applied on each one.
Prompt ordering is a design decision
There is a second effect. Most serving stacks reuse computation for requests sharing a leading prefix. Reuse covers only the span that matches from the start, so putting a variable value near the front invalidates everything behind it.
Prompt ordering therefore carries more than cost and latency — it carries variance. We order ours like this:
The fix, and the trade-off we accepted
Restoring stability means not changing the marker. But the original content-derived scheme keeps the weakness that was flagged. So we tried a middle option: derive the marker per record.
// Seed with the record ID — a value the template author cannot choose in advance
const seed = [...recordId].reduce((a, c) => (a * 31 + c.charCodeAt(0)) | 0, 7);
const fence = [...lines.join('')]
.reduce((a, c) => (a * 31 + c.charCodeAt(0)) | 0, seed) // seed as the initial value
.toString(36);Same record, same marker, so repeated analysis sees identical input. The template author cannot pick the record ID in advance, so the marker cannot be targeted. Measured stability recovered to 2/3.
We did not ship this one either. Deriving the marker per record means the prompt differs between records, which invalidates comparison against every measurement we had already accumulated — we would be rebuilding the evidence base for an accuracy decision. Having confirmed the blast radius, we separated that concern from this change.
The transferable lesson: a security fix can break the very thing the feature exists to do. Type checks passed. Unit tests passed. Every automated gate was green. Only re-running the same audio surfaced it.
A fence is also only one way to isolate untrusted input. The same goal can be served by passing staff-authored text as a field of a structured request rather than concatenating it into the instructions, and by constraining the response with an output schema so that no other shape can come back. The less a system leans on the fence alone, the less a change to the fence can cost.
Was the instruction ignored, or never delivered?
The same investigation surfaced something else.
Our system layers a shared instruction block (applied to every record type) under a format-specific block. The shared block asked for sufficiently detailed prose. The format-specific block asked for one item per line.
When both were present, the model followed the length instruction and ignored the formatting instruction.
| Check | Before | After |
|---|---|---|
| Announcements written one per line | 10 / 18 | 18 / 18 |
| On-call date placed on its own line | 13 / 18 | 18 / 18 |
Precedence between instructions is decided implicitly by the model. It is not reliably “last wins” or “most specific wins.” We resolved it by declaring precedence inside the narrower block — the format-specific one. Editing the shared block would have affected every other record type.
There is a wider rule here: when an instruction fails twice, stop rewriting it and ask whether it can be obeyed at all.
On another feature, a prompt line saying “leave the field empty when nothing applies” was never honoured. The wording was not the problem. The output schema declared that field required, and a required field cannot be empty. No amount of rewriting reaches past a structural constraint.
Our triage order is now:
- Log the exact string that was sent — read what went out, not what you meant to send
- Check whether the output constraints (schema, required fields, minimum lengths) contradict the instruction
- Only if 1 and 2 are clean, edit the wording
The tooling behind the numbers
None of the above is sayable without a measurement rig. Ours has three parts.
First, run the same input repeatedly and take agreement. The implementation is this small:
/** Did every run of the same input produce the same routing? */
const isStable = (runs: Record<string, string[]>[]): boolean =>
runs.every((run) => JSON.stringify(run) === JSON.stringify(runs[0]));Second, pin the denominator. As above, a runner that silently drops outputs inflates the pass rate. Keep the expected number of checks as a constant and fail the run when it does not match.
Third, verify the scorer itself. Whenever we add a guard, we inject a mutation that should trip it and confirm the check turns red. We did this for all four guards added in this change.
The same logic applies to instrumentation. A probe that returns the same value with and without the mutation is measuring nothing. While reproducing one review finding, our first probe reported “no call” whether or not the guard was bypassed — we were one step from reporting the finding as unreproducible. A different probe showed the call plainly and confirmed the finding was real.
What the rig looks like in production
| Layer | What it does | What makes it fail |
|---|---|---|
| Runner | Freezes the transcript, re-runs analysis only, N times | Outputs dropped, shrinking the denominator |
| Scorer | Judges 26 checks in three classes; counts stability per recording | Check count differs from the constant |
| Guard tests | Injects one mutation per guard | A mutated input still passes |
| Post-release | Counts the share of records staff edited by hand each week, always with the denominator | Denominator stays at zero across a week |
That last layer is the one that keeps paying. We describe it in Measuring Medical AI Accuracy by What Clinicians Rewrite. Scoring only before release leaves you blind to the side effects of the next change.
Finally, identical conditions still produce different scores. Our scoring varied by ±4 out of 468. We never decide on a single run; every candidate is measured at least twice. The configuration we shipped measured 454 of 468 checks (97.0%) with all six recordings stable.
What to hold fixed when comparing
| What you want to compare | What must stay fixed |
|---|---|
| Models | Prompt, template, transcript, temperature, thinking budget |
| Templates | Model, transcript |
| Implementations | All of the above. Not one character of the prompt |
We violated this once and had to redo a comparison. We lined up a new model against an old one while the template was two revisions apart, so what looked like a model difference was a template difference. With conditions equalised, the ranking reversed.
The practices, condensed
| Situation | Practice |
|---|---|
| Assembling a prompt | Order it invariant, semi-invariant, variable. Never put a changing value first |
| Passing user-authored text | Send it as a structured field, not concatenated into the instructions |
| An instruction is ignored | Before rewriting it, check the sent string and the output constraints |
| Classification drifts | Read the top-two margin at the deciding token. A thin margin will keep drifting |
| Comparing accuracy | Fix everything except the variable under test, and measure twice |
| Automating scoring | Pin the number of checks; fail the run if it drops |
| Adding a guard | Break it on purpose and confirm the check turns red |
Summary
MENTRA builds structured clinical records for psychiatry from audio. The meeting-minutes workflow described here is part of conference record generation. If your institution already has a fixed minutes format, we assess whether it can be supported as-is — get in touch with the format in hand.
Frequently asked questions
Does temperature 0 make the output identical every time?
No. Temperature 0 takes the highest-probability next token; it does not promise the probabilities are computed identically twice. GPUs sum many parallel results, and that order shifts with concurrent load. Where the top two candidates are close, the shift flips the ranking.
Does pinning a seed make runs reproducible?
Only partially. A seed fixes the sampling draw, but at temperature 0 there is no draw to fix, and the numeric error in computing the scores remains. What measurably worked for us was changing not one character of the prompt, and running the same input several times.
If small prompt edits change the result, is the prompt badly written?
It is separate from writing quality. The model receives a token sequence, not prose. Change 16 meaningless characters and the sequence changes, so the arithmetic changes. Serving stacks also reuse computation for shared prefixes: put invariant text first, variable text last.
