Free · 10 questions · no sign-up · nothing leaves your browser

Eval Readiness Check: is your RAG feature gated, partial, or vibes-only?

Ten questions about how you find out that your RAG or LLM feature answered wrong — and what stops it from happening again. Each answer is worth 0, 1 or 2 points. You get a score out of 20, one of three buckets, and three concrete next steps.

Privacy: there is no backend and no email field. Your answers are kept only in this browser’s localStorage so you can come back to them; “Start over” clears them. No cookies, no analytics, no external requests apart from the fonts.

0–7 · Vibes-only
Changes are checked by hand with a few queries. Every fix is a coin flip.
8–14 · Partial
Some traces, some tests, but nothing blocks a bad release.
15–20 · Gated
A written gate blocks releases. The question is whether it can be trusted.
Question 1Trace capture
For a production request, what do you keep?
Question 2Trace lookup
A user says “it answered wrong”. How fast can you see the exact trace?
Question 3Failure taxonomy
Do you have a written list of failure classes with examples from your own traces?
Question 4Retrieval vs generation
For your last ten wrong answers: was the right document in the context or not?
Question 5Golden set
Do you have a fixed set of queries with expected facts and expected document ids?
Question 6Retrieval metrics
Do you measure retrieval on its own — recall@k and MRR against expected document ids?
Question 7Code checks
Do automated checks run on outputs — citations exist, schema is valid, refusal rules, length?
Question 8Judge calibration
If an LLM judge scores your outputs, has it been compared with human labels?
Question 9Baseline diff
Before changing a prompt, model, chunking or index, do you compare against a stored baseline run?
Question 10Release gate
Is there a written threshold that blocks a release?
0 of 10 answered
Your result
0 / 20
—

Three next steps

    See the kit — pre-order $29 Copied to clipboard