Building an evaluation set that shows an LLM application is ready to release
Public benchmarks rank models; they cannot tell you whether your application answers your users correctly. An evaluation set can: a versioned collection of real and deliberately difficult cases, each with a reference answer or rubric that reviewers agree on. This guide covers choosing what one case represents, sampling cases, scoring them automatically without fooling yourself, and using the set as a release gate.
On this page
- Why public benchmarks cannot approve your release
- Choosing what one test case represents
- Six kinds of case a useful set contains
- Building the first version of the set
- Automated scoring methods and when to trust them
- Calibrating an LLM judge against human reviewers
- Using the set as a release gate
- A hypothetical IT help desk assistant before its pilot
- Keeping an evaluation set honest over time
- Questions and answers
- Sources
Why public benchmarks cannot approve your release
Leaderboards measure general ability on shared tasks. They never see your documents, retrieval settings, instructions, output format or the action that follows. A model that ranks well can still cite a superseded policy, drop a required field or answer a question it should have declined.
An evaluation set turns the release decision into evidence. The NIST AI Risk Management Framework makes measurement one of its four core functions and expects performance to be demonstrated under conditions similar to the deployment setting1. NIST's companion profile for generative AI also lists confabulation, the confident statement of false content, among the risks specific to these systems2. Both points lead to the same practice: cases drawn from your own work, scored in a way you trust and rerun every time something changes.
Choosing what one test case represents
Decide the unit before collecting cases, because it determines what a reference answer looks like and how much each case costs to review.
| Question | Single request | Conversation | End-to-end workflow |
|---|---|---|---|
| What is scored | One input and one output | A sequence of turns, including clarifying questions | The final state of the task, including any action taken in another system |
| Good fit for | Classification, extraction and single-question answering | Assistants where users refine their request over several turns | Agents and assistants that create records, send messages or update systems |
| What it misses | Failures that appear only once context builds up | Whether the outcome was correctly recorded anywhere | Little, although failures are harder to trace to a single step |
| Review effort per case | Low | Medium: reviewers read the whole exchange | High: reviewers check the target system or a simulated copy of it |
Six kinds of case a useful set contains
Sample deliberately: a set built only from typical requests passes applications that fail on the cases that cause incidents.
- Common cases
- The requests that make up most real traffic, sampled from logs or tickets so the mix reflects actual use rather than what the team happens to remember.
- Edge cases
- Legitimate but unusual inputs: very long documents, mixed languages, tables, scanned pages or rarely used policy clauses.
- Missing-information cases
- Requests that cannot be answered from the available sources. The correct behavior is to say so or ask for the missing detail, not to guess.
- Out-of-scope cases
- Requests the application should decline or hand to a person, such as legal advice put to a support assistant.
- Adversarial cases
- Inputs that try to override instructions, extract hidden prompts or data, or smuggle instructions in through a retrieved document.
- Regression cases
- Failures found in testing or production, kept permanently so a fixed problem cannot quietly return.
Building the first version of the set
Write down the release decision
State what the set must prove, for example that answers are correct and cited, that the application declines when it should and that output always fits the downstream schema. Each statement becomes a scoring criterion.
Collect real inputs
Pull anonymized requests from the channels the application will serve. Where none exist yet, ask the people who do the work today to write down the questions they handle, in their own words.
Stratify and fill the gaps
Tag every case by kind and topic, then write new cases for the thin categories, which are usually missing-information and adversarial ones.
Write references and rubrics
For factual tasks, record the correct answer and the source that supports it. For open-ended tasks, write a short rubric: what must be present, what must not appear and what makes an answer fail outright.
Check that reviewers agree
Have two people score the same subset independently. Where they disagree, the rubric is ambiguous; rewrite it before any automated scoring relies on it.
Freeze and version the set
Give the set a version number, store it apart from the prompt repository and record which application version was scored against which set version.
Automated scoring methods and when to trust them
| Method | Works well for | Breaks down when | Check before relying on it |
|---|---|---|---|
| Exact or normalized match | Labels, extracted fields and yes-or-no decisions | Several phrasings are equally correct | That normalization of case, spacing and dates matches what downstream systems accept |
| Schema and structure checks | JSON output, required fields and allowed values | The structure is valid but the content is wrong | That a content check sits alongside it, because structure alone proves little |
| Retrieval and citation checks | Grounded answers that must cite a source | The right passage is cited but misread | That cited passages exist, are current and actually support the claim |
| LLM-as-judge with a rubric | Open-ended answers, completeness and tone | The judge shares the generator's blind spots or favors longer answers | Agreement with human reviewers on a held-out subset |
| Human review | High-stakes or ambiguous cases and rubric design | Volume grows faster than reviewer time | Agreement between reviewers, and fatigue over long sessions |
Calibrating an LLM judge against human reviewers
An automated judge is one model scoring another, so it needs an evaluation of its own. Take a subset of cases that people have already scored with the rubric, run the judge on the same outputs and compare verdicts case by case. Start with the disagreements where the judge passed an answer a person failed, because those are the errors that would let a regression through.
Agreement statistics such as Cohen's kappa summarize the comparison, but reading the disagreements teaches more. Typical fixes are a sharper rubric, one criterion per judging call, asking the judge to quote its evidence and using a different model family from the one under test. Recheck the judge whenever its own model version changes.
After calibration, keep people involved at a lower volume: sample a share of judged cases in each cycle, and send every case the judge marks as uncertain to a reviewer.
Using the set as a release gate
Agree the rules before results arrive, so a disappointing score cannot be argued away afterwards.
- If
A prompt, instruction or tool description changes
ThenRerun the full request-level set and compare each category with the last accepted version.
Small wording changes often fix one category and quietly break another.
- If
The model or model version changes
ThenRerun everything, including workflow-level cases, and recheck the judge's agreement with human scores.
- If
Source documents are re-indexed or retrieval settings change
ThenRun the retrieval and citation cases first, then the full set before release.
- If
A failure is reported in production
ThenAdd it as a regression case, confirm the current version fails it, fix the cause and require the case to pass from then on.
- If
Critical categories pass but the overall score dips
ThenRelease only if the drop is understood and accepted in writing by the product owner; otherwise hold.
A release gate is a decision rule agreed in advance, not a single number.
A hypothetical IT help desk assistant before its pilot
Keeping an evaluation set honest over time
Test cases leak into prompts
Early signalScores jump after failing examples are pasted into the instructions.
MitigationKeep the set in a separate, access-controlled store and draw few-shot examples from a different pool.
The set drifts away from real traffic
Early signalProduction complaints involve topics the set barely covers.
MitigationRe-sample from recent logs on a schedule and retire cases for features that no longer exist.
References go stale
Early signalAnswers that are now correct fail because a policy or product changed.
MitigationLink every reference to its source document and review affected cases when that document changes.
Optimizing for the judge
Early signalJudge scores rise while human spot checks stay flat.
MitigationKeep a human-scored holdout that is never used for tuning, and compare against it at each release.
Questions and answers
How many test cases does an LLM evaluation set need to start?
Enough that every category you care about has several cases, which for a focused application usually means dozens rather than thousands. Coverage of the risky categories matters more than volume. Grow the set from production failures and new features; a small set rerun on every change is worth more than a large one nobody maintains.
Who should own an LLM evaluation set?
The product owner owns the release criteria and signs off the gate. Domain experts own references and rubrics, because only they can say what a correct answer is. Engineering owns the tooling, versioning and the record of which application version passed which set version. Without a named business owner, sets drift toward whatever is easy to score.
Can we generate evaluation cases with an LLM?
Generated cases help fill thin categories, paraphrase real requests and create adversarial variants. They should not replace real inputs, because they tend to be cleaner and more predictable than what users actually write. Have a domain expert review generated cases before they enter the set, and label them so results on real and synthetic cases can be compared.
How often should an LLM application be re-evaluated?
On every change to prompts, tools, retrieval settings or models, and before every release. Fast request-level checks can run automatically in the build pipeline; slower workflow-level and human-reviewed checks run before releases. Add a scheduled rerun too, because provider-side model updates and changing documents can shift behavior with no change on your side.
Sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (Measure function, subcategory 2.3) — National Institute of Standards and Technology · checked 10 October 2026
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 — National Institute of Standards and Technology · checked 10 October 2026