Teardown #1 · synthetic scenario · 9 min read

The refund answer was wrong, and the model was not the problem.

Synthetic scenario. Fjordline Timesheets is a fictional time-tracking product with a fictional help center. Everything below was constructed to show the method on a realistic case; nothing is a measurement from a client system. The method is the one I use on real systems.

A user of Fjordline Timesheets asked the help-center assistant: “I cancelled my annual plan 20 days after paying. Can I get my money back?” The assistant answered, warmly and with a citation, that yes, the first charge is refundable within 30 days of converting to a paid plan, so they were still inside the window.

The policy says 14 days. The user was six days past it. The answer was wrong in the direction that costs money and trust: it promised a refund the company does not give.

The first instinct is to open the prompt and add a line about refunds. I have done that, and each time it fixed the sentence I was looking at and changed one I was not. So the rule this time: nobody touches the prompt until we know how the thing fails.

Thirty traces, sampled, not remembered

A trace is the whole path of one request: the query, the chunks the retriever returned with their document ids, ranks and scores, the prompt version, the answer, the citations, the abstain flag, the latency.

I pulled twenty-five traces from the last two weeks, uniformly at random, and added the refund trace and four others from support tickets. Thirty in all. Uniform first, complaints second, or the taxonomy becomes a list of what people bothered to write in about.

Then, for each trace, one sentence: what went wrong first, upstream. Not the worst thing, the first one. If the retriever missed the right document, that is the sentence; later errors in a trace are mostly shadows of the first.

That first pass is open coding: no categories yet, only sentences like “trial-refund doc ranked above annual-refund doc” and “answer says 5 business days; passage says 5 to 10”. After twenty or so, the sentences repeat and you group them into classes, each with a one-line definition and the stage where it starts.

The pile (synthetic counts):

ClassStageTraces
Similar-but-wrong document ranked firstretrieval5
Right document present, rule split from its exception across chunksretrieval2
Plausible detail added that no passage containsgeneration2
Number present in the passage, bound changed (“within 5 days” for “5 to 10”)generation1
Refused although the context answered the questionpolicy1
Ambiguous query answered without asking which planinput1
No failure found18

Twelve failures out of thirty. Seven of the twelve started in retrieval. In nine of the twelve, the model’s text was faithful to what it had been given.

The refund trace, read properly

The refund query retrieved three chunks. Rank one was the trial policy: if you convert to a paid plan during the trial and change your mind, you can request a refund of that first charge within 30 days of the conversion date. Ranks two and three were the cancellation how-to and the invoices page. The annual-plan refund policy, with its 14-day window, was not in the top five at all.

The model did what it was told: it summarised the retrieved passages and cited them. Every claim in the answer maps to a sentence in the trial document. A faithfulness judge, which checks claims against context, passes this answer; so does a code check that the citations exist. The answer is wrong for the user and right by every check that looks only at the generation step.

Why did the trial document win? Both policies share the words “refund”, “cancel” and “charge”, and the trial chunk is dense with them. The annual-plan page had been split by a fixed-size chunker, so “can be refunded in full within 14 days of the first charge” sat in one chunk and the cancellation sentence in another, each with fewer matching terms. Five of the seven retrieval failures had this shape: two neighbouring policies, the denser one wins.

Fix upstream first

Two changes, both boring:

  1. Chunk by section instead of by size, and prepend the document title and section heading to every chunk, so the text the retriever scores begins with “Refunds for annual plans” or “Trials and trial-to-paid conversion”.
  2. Retrieve five chunks instead of three, so a near miss at rank four still reaches the model.

To know whether it helped, the coded traces became a golden set: for each query, the facts a correct answer must contain, the document ids that count as a hit, and the hard negatives, the similar-but-wrong ids that must not count. Thirty cases. Recall@5 and mean reciprocal rank over that set, old retriever against new (synthetic numbers):

Retrieverrecall@5MRR
Fixed-size chunks, k = 30.630.52
Section chunks with titles, k = 50.870.74

Four cases still miss. Two are the ambiguous-query kind that no retriever can settle; they became “clarification expected” cases. The refund case now has the annual-plan policy at rank one.

Nothing in the prompt changed.

Now the judge

With retrieval measured, the three generation failures were worth a judge. Code cannot tell you that “you can pay by card or bank transfer” was invented; a model reading the context can.

The judge answers one question: is every factual claim in the answer supported by the retrieved context? Two labels, pass or fail. The fail conditions are a numbered list: a claim no passage makes, a claim that contradicts a passage, a rule for one plan stated as universal, a conclusion stitched from two passages that neither states. The procedure asks the judge to list the claims, quote the decisive words of the supporting passage for each one, and only then decide. The output is one JSON object, rationale first, label second.

One criterion per judge. “Is the answer accurate, complete and polite” is three judges, and a combined score cannot be calibrated.

Calibrating it

A judge is a measurement instrument, and it has error rates. Until you know them, its pass rate is a number with an unknown bias.

By day nine there were 100 human-labelled traces for faithfulness: the thirty from the error analysis plus seventy more. A hash of the trace id put each one in a split: 20 train, 40 dev, 40 test. Train is where prompt examples come from, dev is where you iterate, test is touched once per prompt version.

Version one of the judge, on dev, agreed with me on 27 of 28 passes and on 7 of 12 fails. Four of the five misses were the same failure: a number that appears in the passage with a different bound, or a rule for one plan stated for all. The judge saw the number, found it in the passage, and stopped reading. Version two sharpened the fail definitions (a tighter or looser bound counts; a plan-specific rule stated generally counts) and made “quote the decisive words” a required step. On dev: 27 of 28 passes, 11 of 12 fails. Frozen. Then the test split, once:

v2 on the test split (n = 40, synthetic)judge: passjudge: fail
human: pass (30)291
human: fail (10)28

TPR 0.97 (the judge passes what a human passes), TNR 0.80 (it fails what a human fails). Ten fails is thin: about twenty points of uncertainty on that 0.80. The next step is to label more failures, not to admire the number.

On the 60-case regression run after the retrieval fix (the thirty coded cases plus thirty taxonomy-generated variants, tagged synthetic), the judge passed 52: an observed rate of 0.87. Corrected for its error rates, the estimate is also 0.87; here the strictness on good answers and the leniency on bad ones nearly cancel, which is luck, not a law. The bootstrap interval runs from roughly 0.74 to 0.96, wide for the reason above. All three numbers go into the run log next to the judge version.

The gate

The last step is to make the check run without me. A gates file holds the thresholds: recall@5 at least 0.80, MRR at least 0.60, corrected faithfulness pass rate at least 0.85, two must-pass case ids, and zero new failures compared with the stored baseline run. The refund query is a must-pass case; it should never regress quietly.

The harness runs the golden set, writes traces, verdicts and a summary, and the gate command exits 0 for go, 2 for iterate, 1 for block. A GitHub Actions workflow runs it on every change to the prompt, retriever, chunker or corpus, and the report diffs failing cases against the baseline, so a red run asks “which cases newly fail”, not “what is the number”.

The first gated run was a go, recorded as one run-log line: dataset, prompt and index versions, the retrieval numbers, the corrected pass rate with its interval, must-pass failures (zero), decision. The next prompt tweak goes through the same door.

FAQ

Would a faithfulness judge have caught the original wrong answer?
No. The answer was faithful to the wrong document. A judge checks the generation step against what it was given; when retrieval hands over the wrong page, a well-behaved model produces a well-supported wrong answer. Retrieval metrics catch that class, which is why they come first.
Are thirty traces enough?
Enough to find the frequent failure classes and to know where to look first. Not enough to state rates; those come from the harness over the golden set, with an interval, and from the judge’s measured error rates.
Do I need an LLM judge at all?
Only for what code cannot check. Citation exists, JSON is valid, refusal rule fired, answer length: code. Whether a claim is supported by a passage: a judge, calibrated against your own labels before its number reaches a gate.

The files used above, packaged

The worksheet, taxonomy, judge template, calibration sheet, harness and gate are the RAG & LLM Eval Starter Kit. The free sampler on GitHub has the checklist, the judge skeleton and five traces, including the refund one.

See the kit — pre-order $29 Free sampler on GitHub Free 10-question readiness check

Synthetic scenario: Fjordline Timesheets is fictional and every number in this article was constructed for illustration. The method is real.