The refund answer was wrong, and the model was not the problem.
A user of Fjordline Timesheets asked the help-center assistant: “I cancelled my annual plan 20 days after paying. Can I get my money back?” The assistant answered, warmly and with a citation, that yes, the first charge is refundable within 30 days of converting to a paid plan, so they were still inside the window.
The policy says 14 days. The user was six days past it. The answer was wrong in the direction that costs money and trust: it promised a refund the company does not give.
The first instinct is to open the prompt and add a line about refunds. I have done that, and each time it fixed the sentence I was looking at and changed one I was not. So the rule this time: nobody touches the prompt until we know how the thing fails.
Thirty traces, sampled, not remembered
A trace is the whole path of one request: the query, the chunks the retriever returned with their document ids, ranks and scores, the prompt version, the answer, the citations, the abstain flag, the latency.
I pulled twenty-five traces from the last two weeks, uniformly at random, and added the refund trace and four others from support tickets. Thirty in all. Uniform first, complaints second, or the taxonomy becomes a list of what people bothered to write in about.
Then, for each trace, one sentence: what went wrong first, upstream. Not the worst thing, the first one. If the retriever missed the right document, that is the sentence; later errors in a trace are mostly shadows of the first.
That first pass is open coding: no categories yet, only sentences like “trial-refund doc ranked above annual-refund doc” and “answer says 5 business days; passage says 5 to 10”. After twenty or so, the sentences repeat and you group them into classes, each with a one-line definition and the stage where it starts.
The pile (synthetic counts):
| Class | Stage | Traces |
|---|---|---|
| Similar-but-wrong document ranked first | retrieval | 5 |
| Right document present, rule split from its exception across chunks | retrieval | 2 |
| Plausible detail added that no passage contains | generation | 2 |
| Number present in the passage, bound changed (“within 5 days” for “5 to 10”) | generation | 1 |
| Refused although the context answered the question | policy | 1 |
| Ambiguous query answered without asking which plan | input | 1 |
| No failure found | 18 |
Twelve failures out of thirty. Seven of the twelve started in retrieval. In nine of the twelve, the model’s text was faithful to what it had been given.
The refund trace, read properly
The refund query retrieved three chunks. Rank one was the trial policy: if you convert to a paid plan during the trial and change your mind, you can request a refund of that first charge within 30 days of the conversion date. Ranks two and three were the cancellation how-to and the invoices page. The annual-plan refund policy, with its 14-day window, was not in the top five at all.
The model did what it was told: it summarised the retrieved passages and cited them. Every claim in the answer maps to a sentence in the trial document. A faithfulness judge, which checks claims against context, passes this answer; so does a code check that the citations exist. The answer is wrong for the user and right by every check that looks only at the generation step.
Why did the trial document win? Both policies share the words “refund”, “cancel” and “charge”, and the trial chunk is dense with them. The annual-plan page had been split by a fixed-size chunker, so “can be refunded in full within 14 days of the first charge” sat in one chunk and the cancellation sentence in another, each with fewer matching terms. Five of the seven retrieval failures had this shape: two neighbouring policies, the denser one wins.
Fix upstream first
Two changes, both boring:
- Chunk by section instead of by size, and prepend the document title and section heading to every chunk, so the text the retriever scores begins with “Refunds for annual plans” or “Trials and trial-to-paid conversion”.
- Retrieve five chunks instead of three, so a near miss at rank four still reaches the model.
To know whether it helped, the coded traces became a golden set: for each query, the facts a correct answer must contain, the document ids that count as a hit, and the hard negatives, the similar-but-wrong ids that must not count. Thirty cases. Recall@5 and mean reciprocal rank over that set, old retriever against new (synthetic numbers):
| Retriever | recall@5 | MRR |
|---|---|---|
| Fixed-size chunks, k = 3 | 0.63 | 0.52 |
| Section chunks with titles, k = 5 | 0.87 | 0.74 |
Four cases still miss. Two are the ambiguous-query kind that no retriever can settle; they became “clarification expected” cases. The refund case now has the annual-plan policy at rank one.
Nothing in the prompt changed.
Now the judge
With retrieval measured, the three generation failures were worth a judge. Code cannot tell you that “you can pay by card or bank transfer” was invented; a model reading the context can.
The judge answers one question: is every factual claim in the answer supported by the retrieved context? Two labels, pass or fail. The fail conditions are a numbered list: a claim no passage makes, a claim that contradicts a passage, a rule for one plan stated as universal, a conclusion stitched from two passages that neither states. The procedure asks the judge to list the claims, quote the decisive words of the supporting passage for each one, and only then decide. The output is one JSON object, rationale first, label second.
One criterion per judge. “Is the answer accurate, complete and polite” is three judges, and a combined score cannot be calibrated.
Calibrating it
A judge is a measurement instrument, and it has error rates. Until you know them, its pass rate is a number with an unknown bias.
By day nine there were 100 human-labelled traces for faithfulness: the thirty from the error analysis plus seventy more. A hash of the trace id put each one in a split: 20 train, 40 dev, 40 test. Train is where prompt examples come from, dev is where you iterate, test is touched once per prompt version.
Version one of the judge, on dev, agreed with me on 27 of 28 passes and on 7 of 12 fails. Four of the five misses were the same failure: a number that appears in the passage with a different bound, or a rule for one plan stated for all. The judge saw the number, found it in the passage, and stopped reading. Version two sharpened the fail definitions (a tighter or looser bound counts; a plan-specific rule stated generally counts) and made “quote the decisive words” a required step. On dev: 27 of 28 passes, 11 of 12 fails. Frozen. Then the test split, once:
| v2 on the test split (n = 40, synthetic) | judge: pass | judge: fail |
|---|---|---|
| human: pass (30) | 29 | 1 |
| human: fail (10) | 2 | 8 |
TPR 0.97 (the judge passes what a human passes), TNR 0.80 (it fails what a human fails). Ten fails is thin: about twenty points of uncertainty on that 0.80. The next step is to label more failures, not to admire the number.
On the 60-case regression run after the retrieval fix (the thirty coded cases plus thirty taxonomy-generated variants, tagged synthetic), the judge passed 52: an observed rate of 0.87. Corrected for its error rates, the estimate is also 0.87; here the strictness on good answers and the leniency on bad ones nearly cancel, which is luck, not a law. The bootstrap interval runs from roughly 0.74 to 0.96, wide for the reason above. All three numbers go into the run log next to the judge version.
The gate
The last step is to make the check run without me. A gates file holds the thresholds: recall@5 at least 0.80, MRR at least 0.60, corrected faithfulness pass rate at least 0.85, two must-pass case ids, and zero new failures compared with the stored baseline run. The refund query is a must-pass case; it should never regress quietly.
The harness runs the golden set, writes traces, verdicts and a summary, and the gate command exits 0 for go, 2 for iterate, 1 for block. A GitHub Actions workflow runs it on every change to the prompt, retriever, chunker or corpus, and the report diffs failing cases against the baseline, so a red run asks “which cases newly fail”, not “what is the number”.
The first gated run was a go, recorded as one run-log line: dataset, prompt and index versions, the retrieval numbers, the corrected pass rate with its interval, must-pass failures (zero), decision. The next prompt tweak goes through the same door.
FAQ
Would a faithfulness judge have caught the original wrong answer?
Are thirty traces enough?
Do I need an LLM judge at all?
The files used above, packaged
The worksheet, taxonomy, judge template, calibration sheet, harness and gate are the RAG & LLM Eval Starter Kit. The free sampler on GitHub has the checklist, the judge skeleton and five traces, including the refund one.
See the kit — pre-order $29 Free sampler on GitHub Free 10-question readiness checkSynthetic scenario: Fjordline Timesheets is fictional and every number in this article was constructed for illustration. The method is real.