Skip to content
NakodaAI

AI FIELD MANUALDEVELOPERS, FOUNDERS

The demo worked. Before this LLM feature ships, how do you know it'll still work after the next prompt change?

A real evaluation methodology - golden dataset, LLM-as-judge with a written rubric, and a regression gate - for teams about to ship an AI feature, instead of shipping on "it looked right in testing."

Last reviewed 1 September 2026

THE PROBLEM

A small product team has built a customer-support triage feature: an LLM reads an incoming ticket and classifies it into one of eight categories, then drafts a first response. It works well on the dozen tickets the developer tried during the demo. There is no test suite for it, because "testing an AI feature" doesn't obviously map to the unit tests the rest of the codebase has.

Two weeks after launch, someone tweaks the system prompt to fix a specific misclassification they noticed, ships it same-day, and three other categories quietly get worse. Nobody notices until a customer complains, because there was nothing in place that would have caught the regression before it reached production.

The team needs the AI equivalent of a test suite - something that runs before a prompt or model change ships and answers "did this change make things better or worse", not "did it fix the one example I was looking at."

THE APPROACH

Build a golden dataset - a fixed, representative set of real (or realistic) inputs with the correct expected output for each, written and agreed on by a human before any evaluation runs. Then use an LLM-as-judge - a second model call that scores each output against a written rubric - to score the feature's outputs at scale, and gate any prompt or model change on the aggregate score not regressing.

This mirrors how a conventional test suite works: fixed inputs, known-correct expectations, run automatically before every change ships. The only new part is that grading free-text or classification output at scale needs a model to do the grading, because a human can't manually review every regression run - which is exactly why the rubric the judge model uses has to be written down and version-controlled like code, not left implicit.

WHY IT WORKS

A golden dataset turns "does this look better" (a vibe, subject to whoever is looking) into "does this score better against 50 fixed, agreed-on examples" (a number, comparable run over run). This is the entire reason software testing exists, applied to a system whose outputs are less deterministic than a function's return value.

LLM-as-judge works because grading against a specific written rubric ("is the category correct: yes/no. Is the drafted response polite and on-topic: 1-5") is a much narrower, more reliable task for a model than the original open-ended generation task - judging is closer to classification than to creative generation, and narrower tasks are where LLMs are most consistent.

The regression gate is what actually prevents the failure in the problem statement: a change that improves the aggregate score on the golden set can ship; a change that drops it cannot, automatically, before a human ever has to notice in production.

STEP BY STEP

  1. 01

    Golden dataset

    40-100 real examples with human-written correct answers, fixed and version-controlled.

  2. 02

    Candidate run

    The actual feature pipeline run against every golden example.

  3. 03

    LLM-as-judge

    A second model scores each output against a written, explicit rubric - not a vague quality check.

  4. 04

    Aggregate score

    One comparable number (or per-category breakdown) for this run.

  5. 05

    Regression gate

    Change ships only if the score doesn't regress - automated, not a vibe check.

  1. Collect 40-100 real or realistic examples

    Pull actual historical tickets if you have them (redacted of customer PII), or write realistic ones covering every category and the edge cases you already know are hard. Fewer than ~40 examples makes the aggregate score noisy; more than a few hundred is usually unnecessary for a first version.

  2. Write the correct expected output for each, by hand

    A human decides and writes down the correct category and an example of an acceptable response for every single example. This step cannot be delegated to the model being evaluated - that would be grading its own homework.

  3. Write the judge rubric explicitly

    For each dimension being scored (category correctness, tone, factual accuracy, whether it asked for missing information appropriately), write the exact instruction the judge model will be given, with a defined scale. "Rate quality 1-5" without criteria produces an unreliable judge; "Rate 1-5: 5 = correctly identifies the issue and asks one clarifying question if information is missing, 1 = misidentifies the issue" produces a usable one.

  4. Run the feature against the golden set

    Automate running every golden example through the actual feature pipeline (same prompt, same model, same post-processing the real feature uses) and capturing the output.

  5. Run the judge model against every output

    For each output, call a judge model with the rubric, the golden expected answer, and the actual output, and record the score. Spot-check a sample of the judge's scores by hand early on to confirm the judge itself is behaving as intended.

  6. Set a gate and wire it into the ship process

    Define a minimum aggregate score (or: no category may regress by more than X%) that a prompt or model change must meet before it can ship. Run the suite automatically on every change, the same way a CI test suite runs on every pull request.

TOOLS

LIMITATIONS

  • A golden set of 50-100 examples cannot cover every real-world input; it catches regressions on the patterns you already thought to include, not novel failure modes. Add real production failures to the golden set as they're discovered, so it grows over time rather than staying frozen at launch.

  • LLM-as-judge inherits its own biases - it can be measurably swayed by response length, formatting, or which model produced the answer, independent of actual quality. Spot-checking the judge against human judgment periodically (not just once at setup) is not optional.

  • This catches regressions in what you decided to measure. If tone matters to customers but isn't in the rubric, the suite will happily pass a change that gets colder and less helpful in a way no dimension was scoring for.

EXAMPLE

A support-ticket triage feature needs a regression gate before the team ships a system-prompt fix for one misclassified category.

  1. 60 historical tickets collected across all 8 categories, redacted, with a human-assigned correct category and one example acceptable response each.

  2. Rubric written: category correctness (pass/fail), tone (1-5, defined scale), whether a clarifying question was asked when the ticket was ambiguous (pass/fail).

  3. Baseline run: 91% category accuracy, average tone 4.2, clarifying-question rate 68%.

  4. Prompt changed to fix the specific misclassification. Re-run: category accuracy rose to 94%, but clarifying-question rate dropped to 51% - the new prompt wording had made the model more decisive and less likely to ask when it should have.

  5. The regression on the third dimension was caught before shipping; the prompt was revised again to fix the original issue without losing the clarifying-question behavior, and the suite re-run before merge.

A regression that would have shipped invisibly - fewer clarifying questions, more confidently wrong first responses - caught by the gate before a single real customer saw it.

RELATED

Nakoda editorial · last reviewed

This entry describes a workflow Nakoda recommends - it is not a claim about how any named tool behaves in every case, and it is not paid placement. Spotted something out of date? Tell us.

MORE FOR THIS AUDIENCE