RAG & LLM Eval Starter Kit · Core V1 · Pre-order

From “it sometimes answers wrong” to a repeatable release gate — in one afternoon.

RAG & LLM Eval Starter Kit: for developers whose shipped RAG or LLM feature “sometimes answers wrong” — a downloadable kit that turns your 30 worst traces into a failure taxonomy, a golden test set and a Python regression harness in 14 days at under an hour a day. No platform, no calls, no subscription. 14-day refund if it does not fit your stack.

The first gated run on your own pipeline takes an afternoon: the harness runs on the bundled example with no API key, and your predict() adapter is one function. Making the gate trustworthy — a taxonomy from your traces, a calibrated judge, a golden set — is what the 14 days are for. No eval platform, no consultant.

Pre-orders open Oct 7 — $29 (ships by Oct 21) Free 10-question readiness check
Python ≥ 3.10 · stdlib only Provider-agnostic ZIP + private GitHub release Single-team license
python -m evalkit gate 30 tracestaxonomygolden setharness + judgegate traces → taxonomy → golden set → harness → gate

Who it is for — and who it is not for

Read this before the price

For you ifSomething is in production and it sometimes answers wrong

  • You have shipped a RAG or LLM feature — document Q&A, a support assistant, an internal helper — and users or teammates report wrong answers.
  • You are a team of one to five. Today you change the model, the prompt or the chunking and check by hand with a few queries.
  • A demo, a first customer or a model migration is coming, and you need to know what will break before it does.
  • You are comfortable in Python and a terminal, and you want files you own, not another dashboard.

Not for you ifYou already have the machinery, or nothing to measure yet

  • Your team runs an eval platform with someone whose job is evals. The kit would duplicate what you have.
  • Nothing is in production yet, so there are no traces to sample. Take the readiness check and come back when you have thirty bad ones.
  • You want video lessons, a community or live calls. This is files and written support only.
  • You need multi-turn agent or tool-use evaluation. Core V1 covers single-turn RAG and LLM features.

What is inside Core V1

Every file is original · examples are synthetic
00
Eval plan template
One page that pins down what you measure, the single North-Star metric, and the threshold chosen by the cost of a wrong answer — before any tooling.
kit/00-eval-plan.md
M1
Error-analysis protocol + taxonomy
A sampling plan (uniform first, then stratified), a coding worksheet with a “first upstream error” column, an eight-class failure taxonomy, a saturation rule that tells you when to stop reading, and a RAG watchlist of what to look at in each trace.
kit/01-error-analysis/ — protocol.md · taxonomy.csv · worksheet.csv · sampling-plan.md · rag-watchlist.md · saturation-rule.md
M2
Judge skeleton + calibration sheet
A judge prompt skeleton (task, pass/fail definitions, hard examples, rationale before verdict, strict JSON), three ready RAG judges — faithfulness, context-use, refusal-appropriateness — a train/dev/test calibration sheet with TPR/TNR by segment, corrected pass rate with a confidence interval, and a “judge failure zoo” of ways judges go wrong.
kit/02-judges/ — judge-skeleton.md · calibration.md · calibration-sheet.csv · judge-failure-zoo.md · judges/*.md
M3
Retrieval + generation sheets
A retrieval sheet (expected ids, top-k, first hit rank, hit@1/3/5, reciprocal rank), a claim-level generation sheet (supported / unsupported / contradicted; citation exists, citation supports), a corpus health checklist, a synthetic question generator with hard negatives, and a retrieval error-analysis guide.
kit/03-retrieval/ — retrieval-sheet.csv · generation-claim-sheet.csv · corpus-health-checklist.md · synthetic-question-generator.md · retrieval-error-analysis.md
M4
Runnable Python harness + CI regression gate
The evalkit package: JSONL dataset loader, pure-Python BM25 retriever, retrieval metrics, LLM providers (deterministic mock, any OpenAI-compatible endpoint, Anthropic — via urllib, no SDKs) with a hash-keyed cache, judge runner, calibration maths, run store, markdown report and a gate command that exits 0 / 1 / 2 for go / block / iterate. Plus a GitHub Actions workflow that runs it on every change. Standard library only, Python 3.10+.
evalkit/ · examples/ · tests/ · .github/workflows/regression.yml · gates.json
M5
14-day rollout plan + release gate
The day-by-day plan below as a checklist, a release-gate checklist, and a run log with one row per run: dataset version, model, prompt version, index version, recall@5, MRR, corrected judge pass rate with interval, must-pass failures, decision.
kit/05-rollout/ — 14-day-plan.md · release-gate-checklist.md · run-log.csv
M6
Synthetic traces by dimensions
For when you have too few real traces: a dimensions worksheet, a realism rubric to prune what would never happen, provenance tags (synthetic / hand / prod) so the two never mix silently, and a generator prompt skeleton.
kit/06-synthetic-traces/ — dimensions-worksheet.md · realism-rubric.md · provenance.md · generator-prompt-skeleton.md
M7
Minimal annotation reviewer + protocol
A single-file Python reviewer (CSV in, CSV out, hotkeys, autosave) for labelling traces without a web app, and a short annotation protocol: who labels, how much overlap, how disagreements are resolved, and how to read Cohen’s kappa.
kit/07-annotation/ — reviewer.py · annotation-protocol.md
EX
Worked example: Fjordline Timesheets
A fictional SaaS help center as a toy corpus, a golden set of 20 synthetic cases in the kit’s case format, a toy pipeline and a one-command run that produces a run directory and a report with no network and no key. Labelled synthetic throughout.
examples/toy-corpus/ · examples/golden.jsonl · examples/toy_pipeline.py · examples/run_toy.py
Formats are fixed and documented (case JSONL, trace JSON, judge verdict JSON, CSV headers) so the sheets, the harness and the reviewer all read the same files. Nothing in the kit phones home.

How the 14 days work

Under an hour a day · the plan file has the checklists

Pricing

One-time payment · no subscription · checkout by Lemon Squeezy
Written Eval Audit
$249
Written Eval Audit — $249

You send 30 traces and your configuration (chunking, retrieval, prompt). Within 5 business days you get a written failure taxonomy built from your traces and a prioritised list of fixes. Written only, no calls.

  • You send: 30 redacted traces in the kit’s trace format, plus retrieval and prompt config.
  • You get: a written diagnosis — failure classes with counts, retrieval versus generation split, prioritised fixes, suggested gate thresholds.
  • Timing: intake reply within 24 hours; report within 5 business days of receiving the traces.
  • Not included: implementation, calls, ongoing monitoring.
Written audit intake opens Oct 7
Fixed scope, fixed price. If your traces show the audit will not help (for example, there is no retrieval step), I say so before you pay.
The kit · Core V1
$29$49
Pre-order $29 — ships by Oct 21, 2026 (then $49)

Everything listed above: modules 00–M7, the evalkit harness with CI workflow, the synthetic worked example and tests. Delivered as a ZIP download and a private GitHub release.

  • Single-team license — use it on any number of your own projects.
  • 14-day refund from delivery, no questions asked.
  • Written support by email, answered within 24 hours.
  • Core V1 updates included; no subscription, ever.
Pre-orders open Oct 7 — $29 (ships by Oct 21)
Pre-order price holds until Oct 21, 2026 or the first 30 orders, whichever comes first. If the kit is ready earlier, it ships earlier. Lemon Squeezy handles checkout, VAT and the receipt; you get the download email when it ships.

This is a V1. The module list above is the exact scope; the FAQ says what is not included. There are no testimonials on this page because the kit has not shipped yet.

Questions

Written support: reply within 24 h
Why not use a free eval library?

Use one — the harness sits next to them, it does not replace them. Free tools give you an API: a metric function, a judge class, a tracing UI. They do not tell you which thirty traces to read, how to code them, when to stop, whether your judge agrees with a human, or what number should block a release.

The kit is that workflow, as files with fixed columns and a day-by-day order. The harness is standard-library Python so it runs anywhere your code runs, and the trace format is plain JSON, so you can export from whatever you already use.

What do I get compared with a course?

Files, not video. Worksheets with exact columns, judge prompts you paste and run, a harness with tests, a plan with a checklist per day. A course explains the idea; the kit assumes you understand the idea and need the artefacts by Day 14.

There is no theory chapter. Where a method needs explaining, the file has a two-to-four line header saying what it is for and when to use it.

When does it ship, and what is the refund rule?

Pre-orders ship by Oct 21, 2026 as a ZIP download plus access to a private GitHub release. If it is ready earlier, it ships earlier and you get an email the same day.

Refund: 14 days from delivery, no questions asked. Reply to the delivery email with the word “refund” and Lemon Squeezy returns the payment to the original method. If you pre-order and change your mind before delivery, the same applies.

What is not included?

No hosted platform or dashboard. No video lessons, community or live calls. No API credits — judges on your own traces use your own key. No integration code for specific vector databases or frameworks; you write one predict() function and the harness calls it. No multi-turn agent or tool-use evaluation in V1. No promise of any particular accuracy number — the kit measures, it does not guarantee.

What does “single-team license” mean?

One purchase covers one team: the people working on one product or codebase. Use the files on any number of your own projects, modify them freely, keep the harness in your private repos. You may not redistribute or resell the kit files, or repackage them as a course or a product. The free public sampler is MIT and separate from this license.

Do I need API keys?

Not to run the kit. The bundled example runs on a deterministic mock LLM with no network. Retrieval metrics, code checks, calibration maths and the gate never call a model.

To run LLM judges on your own traces you need one model endpoint: any OpenAI-compatible URL (including a local server) or Anthropic, with your own key in an environment variable. Judge calls are cached by hash, so re-runs are free.

How does support work?

Written only, by email, answered within 24 hours. Send the command you ran, the error, or the trace you are stuck on. No calls, no screen shares. If a question turns out to be a bug in the kit, the fix ships to everyone.

Are the examples real data?

No. Every example — the Fjordline Timesheets help center, the golden set, the traces in the judge zoo — is synthetic, written for this kit and labelled as such. No client data, no scraped content. The methods were shaped by production work; the data was not taken from it.

A note from the author

I build retrieval-augmented assistants that people rely on at work — document search over maintenance manuals, internal knowledge tools, support-style question answering — and I write the evaluations that decide whether a change ships. Every one of those systems produced a “that answer was wrong” message eventually. The useful question was never “which model”. It was: was the right chunk even in the context, and how would I know next time?

This kit is the set of files I wish I had had when the first such message arrived: the worksheet, the judge prompt, the calibration sheet and a harness small enough to read in an evening. It is a V1 from one practitioner, not a platform. What it promises is on this page, and nothing more.

— Laimonas Breidokas · AuditLab, independent AI lab
Principles the kit follows
Read traces before touching promptsFix upstream firstCode checks before judgesCalibrate before trustingGate every changeSynthetic is labelled