RAG & LLM Eval Starter Kit: for developers whose shipped RAG or LLM feature “sometimes answers wrong” — a downloadable kit that turns your 30 worst traces into a failure taxonomy, a golden test set and a Python regression harness in 14 days at under an hour a day. No platform, no calls, no subscription. 14-day refund if it does not fit your stack.
The first gated run on your own pipeline takes an afternoon: the harness runs on the bundled example with no API key, and your predict() adapter is one function. Making the gate trustworthy — a taxonomy from your traces, a calibrated judge, a golden set — is what the 14 days are for. No eval platform, no consultant.
evalkit package: JSONL dataset loader, pure-Python BM25 retriever, retrieval metrics, LLM providers (deterministic mock, any OpenAI-compatible endpoint, Anthropic — via urllib, no SDKs) with a hash-keyed cache, judge runner, calibration maths, run store, markdown report and a gate command that exits 0 / 1 / 2 for go / block / iterate. Plus a GitHub Actions workflow that runs it on every change. Standard library only, Python 3.10+.predict(query) for your own pipeline — answer, citations, retrieved chunks, abstain flag — and make sure production writes the same trace shape.gates.json: retrieval minimums, judge pass minimum, must-pass case ids, zero new failures versus baseline.You send 30 traces and your configuration (chunking, retrieval, prompt). Within 5 business days you get a written failure taxonomy built from your traces and a prioritised list of fixes. Written only, no calls.
Everything listed above: modules 00–M7, the evalkit harness with CI workflow, the synthetic worked example and tests. Delivered as a ZIP download and a private GitHub release.
This is a V1. The module list above is the exact scope; the FAQ says what is not included. There are no testimonials on this page because the kit has not shipped yet.
Use one — the harness sits next to them, it does not replace them. Free tools give you an API: a metric function, a judge class, a tracing UI. They do not tell you which thirty traces to read, how to code them, when to stop, whether your judge agrees with a human, or what number should block a release.
The kit is that workflow, as files with fixed columns and a day-by-day order. The harness is standard-library Python so it runs anywhere your code runs, and the trace format is plain JSON, so you can export from whatever you already use.
Files, not video. Worksheets with exact columns, judge prompts you paste and run, a harness with tests, a plan with a checklist per day. A course explains the idea; the kit assumes you understand the idea and need the artefacts by Day 14.
There is no theory chapter. Where a method needs explaining, the file has a two-to-four line header saying what it is for and when to use it.
Pre-orders ship by Oct 21, 2026 as a ZIP download plus access to a private GitHub release. If it is ready earlier, it ships earlier and you get an email the same day.
Refund: 14 days from delivery, no questions asked. Reply to the delivery email with the word “refund” and Lemon Squeezy returns the payment to the original method. If you pre-order and change your mind before delivery, the same applies.
No hosted platform or dashboard. No video lessons, community or live calls. No API credits — judges on your own traces use your own key. No integration code for specific vector databases or frameworks; you write one predict() function and the harness calls it. No multi-turn agent or tool-use evaluation in V1. No promise of any particular accuracy number — the kit measures, it does not guarantee.
One purchase covers one team: the people working on one product or codebase. Use the files on any number of your own projects, modify them freely, keep the harness in your private repos. You may not redistribute or resell the kit files, or repackage them as a course or a product. The free public sampler is MIT and separate from this license.
Not to run the kit. The bundled example runs on a deterministic mock LLM with no network. Retrieval metrics, code checks, calibration maths and the gate never call a model.
To run LLM judges on your own traces you need one model endpoint: any OpenAI-compatible URL (including a local server) or Anthropic, with your own key in an environment variable. Judge calls are cached by hash, so re-runs are free.
Written only, by email, answered within 24 hours. Send the command you ran, the error, or the trace you are stuck on. No calls, no screen shares. If a question turns out to be a bug in the kit, the fix ships to everyone.
No. Every example — the Fjordline Timesheets help center, the golden set, the traces in the judge zoo — is synthetic, written for this kit and labelled as such. No client data, no scraped content. The methods were shaped by production work; the data was not taken from it.