Founder & Lead Engineer, RAITHub
Test an LLM feature in three layers: ordinary unit tests for the code around the model, a golden set of real inputs scored by code-based checks wherever possible, and a CI gate that fails the build when the pass rate drops below a baseline. Use an LLM as judge only for qualities code cannot check, and calibrate it against human labels first.
An LLM feature is any part of your product where a large language model produces the output: a support answer, a summary, an extracted field, a classification, a tutoring hint. Ordinary tests assume the same input gives the same output. Model output varies from run to run and changes when the prompt or model changes, so the tests have to measure a rate, not a single result. This guide shows how, from the QA side.
Why can't you test an LLM feature like normal code?
You can for most of it. Only the model call itself needs a different approach. OpenAI's evaluation best practices describe evals as "a way to test your AI system despite this variability", and as one of the only ways to improve an LLM application's performance. An eval (short for evaluation) is a repeatable test that runs the feature over a set of inputs and scores the outputs.
Split the feature into parts and test each the way it deserves:
| Part of the feature | Deterministic? | How to test it |
|---|---|---|
| Prompt assembly, retrieval query, tool arguments | Yes | Unit tests with fixed inputs and exact expected values |
| Output parsing and schema validation | Yes | Unit tests, including malformed and empty model responses |
| Guardrails: length limits, refusals, PII filters | Yes | Unit tests with adversarial inputs |
| The model's answer quality | No | A golden-set eval with a pass-rate threshold |
| Behaviour in production | No | Logging, sampled review, and feeding failures back into the golden set |
Most defects in LLM features are in the deterministic parts: a prompt built from the wrong record, a parser that crashes on an extra field, a missing timeout. Those tests are cheap, fast and do not call a model at all. Write them first.
What is a golden set, and how big should it be?
A golden set is a fixed, versioned list of inputs with what a correct output must contain or avoid. It is the LLM feature's regression suite.
- Use real inputs. Anthropic's guide to building evaluations says to mirror your real-world task distribution and include edge cases: irrelevant or missing input data, overly long inputs, poor or harmful user input, and ambiguous cases.
- Prefer volume. The same guide advises more test cases with slightly lower-signal automated grading over fewer hand-graded ones. Start with 50 to 100 cases you can run in minutes; grow from there.
- Write expectations as properties, not exact strings: must mention the refund window, must not promise a delivery date, must return valid JSON with a
categoryfield. - Mine your logs. OpenAI's guide recommends logging as you develop so you can mine the logs for good eval cases. Every production failure becomes a new case.
- Version it with the code. The golden set lives in the repository and changes through review, like tests.
An entry can be as simple as this:
{
"id": "refund-window-01",
"input": "I bought this 20 days ago, can I still return it?",
"mustInclude": ["30 days"],
"mustNotInclude": ["guarantee", "no receipt needed"]
}
How do you grade the outputs?
Use the least expensive grader that is reliable for the property, and move up only when you must. Anthropic's guide lists the options in roughly this order: exact match, string checks, code-graded checks, then LLM-based grading.
| Grader | Good for | Cost per run | Reliability |
|---|---|---|---|
| Exact match | Classification, extracted IDs, yes/no answers | Free | High |
| String and regex checks | Required facts, forbidden phrases, format | Free | High for what they cover |
| Code-based checks | JSON schema, arithmetic, SQL that runs, citations that exist | Free or cheap | High |
| Embedding similarity | Consistency across rephrased questions | Low | Medium; thresholds need tuning |
| LLM-as-judge | Tone, helpfulness, faithfulness to a source | A model call per case | Medium; has known biases |
| Human review | Calibrating the others, high-stakes samples | High | High, but slow |
Code-based checks go further than people expect. PadhAI, the AI tutoring platform RAITHub built, runs a Socratic tutor with math verification: the tutor's mathematical claims are checked rather than trusted. Wherever a model's output can be verified by computation, a query or a lookup, verify it in code instead of asking another model.
What are the caveats of LLM-as-judge?
LLM-as-judge means asking a second model to score the first model's output against a rubric. It is useful for qualities code cannot check, and it has measured weaknesses.
- It agrees with people often, not always. In Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023), strong judges such as GPT-4 reached over 80% agreement with human preferences, about the level humans reach with each other. That still leaves roughly one verdict in five in dispute.
- Position bias. The same paper and OpenAI's guide both report that judges favour an answer because of the order it appears in. In pairwise comparisons, run each pair in both orders.
- Verbosity bias. Judges tend to prefer longer answers. OpenAI's guide suggests controlling for response length.
- Self-enhancement bias. Zheng et al. found judges can favour output from their own model. Anthropic's guide advises using a different model to evaluate than the one being evaluated.
- Weak reasoning. The paper also notes judges' limited ability to grade math and reasoning. Use code for those.
To control these: prefer pass/fail or pairwise questions over 1-to-10 scores, as OpenAI recommends; give the judge a short, specific rubric; pin the judge model version so scores stay comparable; and before trusting it, have a person label 30 to 50 cases and check the judge agrees with them. Recheck that agreement whenever you change the judge.
How do you gate releases on evals in CI?
Run the golden set on every change that can alter model behaviour, and fail the build if the pass rate falls below a threshold. OpenAI's guide puts it plainly: run evals on every change, monitor for new cases of nondeterminism, and grow the eval set over time.
A minimal TypeScript version with Vitest. It uses only code-based grading, so it is deterministic apart from the model call:
// evals/support-answer.eval.test.ts
import { readFileSync } from 'node:fs'
import { describe, expect, it } from 'vitest'
import { answerSupportQuestion } from '../src/ai/support-answer'
interface GoldenCase {
id: string
input: string
mustInclude: string[]
mustNotInclude: string[]
}
const cases: GoldenCase[] = JSON.parse(
readFileSync('evals/golden/support-answer.json', 'utf8'),
)
// Set from a measured baseline, not a guess. Raise it as the feature improves.
const MIN_PASS_RATE = 0.9
function passes(output: string, c: GoldenCase): boolean {
const text = output.toLowerCase()
return (
c.mustInclude.every((s) => text.includes(s.toLowerCase())) &&
c.mustNotInclude.every((s) => !text.includes(s.toLowerCase()))
)
}
describe('support answer golden set', () => {
it('meets the pass-rate gate', async () => {
const failed: string[] = []
for (const c of cases) {
const output = await answerSupportQuestion(c.input)
if (!passes(output, c)) failed.push(c.id)
}
const passRate = (cases.length - failed.length) / cases.length
console.log('pass rate', passRate.toFixed(3), 'failed:', failed.join(', ') || 'none')
expect(passRate).toBeGreaterThanOrEqual(MIN_PASS_RATE)
}, 300_000)
})
Then run it in CI only when prompts, model settings or the AI code change, so ordinary pull requests stay fast and cheap:
# .github/workflows/evals.yml
name: llm-evals
on:
pull_request:
paths: ['src/ai/**', 'prompts/**', 'evals/**']
jobs:
evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 24
- run: npm ci
- run: npx vitest run evals
env:
OPENAI_API_KEY: ${{ secrets.EVAL_OPENAI_API_KEY }}
Four practical rules make the gate trustworthy:
- Set the threshold from a baseline. Run the set several times on the current version, and set the gate a little below the lowest result, so model noise alone does not fail builds.
- Fix model settings. Pin the model version and use a low temperature (the setting that controls randomness) for the eval run, where the provider supports it.
- Use a separate, capped key. Evals cost money on every run; give CI its own key with a spending limit.
- Report the failed IDs, not only the rate, so a reviewer can read the actual outputs.
What should you monitor after release?
The golden set only covers inputs you thought of. In production, log inputs, outputs, model version and latency; review a small sample regularly; track user signals such as thumbs-down or escalations; and turn every confirmed failure into a golden-set case. That loop is what keeps the gate relevant. For the wider design of an LLM feature, including routing, validation and rollout, see how to add AI features to an existing SaaS; for a full AI product, see how to build an AI tutor.
Why RAITHub for this?
Because RAITHub treats LLM features as software to be tested, not magic to be trusted. RAITHub works QA-first: this website carries 400+ tests, and the method is written up in how RAITHub tests software. On the AI side, RAITHub built PadhAI, with 11 services, a 70/20/10 model router and a tutor whose math is verified rather than taken on trust. RAITHub does not publish PadhAI's eval scores, so treat the numbers in this guide as starting points, not results. The QA and test automation service sets up golden sets, graders and CI gates on your stack; for browser-level tests around the feature, see Playwright vs Cypress.
When should you use a tool instead?
- You want a ready-made eval runner. Open-source and hosted eval tools can store datasets, run graders and show dashboards. If your team can own the golden set and thresholds, a tool may be all you need.
- The feature is low-stakes. A draft suggestion a person always edits may need only unit tests around the model and a small sampled review.
- You have a QA engineer already running CI. This guide and the code sample are enough to start.
Last reviewed: 29 September 2026.
If your LLM feature is shipping without a regression gate, ask RAITHub about QA for AI features. The free 15-minute technical audit looks at what you test today, and you get a fixed written quote for the golden set, graders and CI gate.
Frequently asked questions
How do you test an LLM feature?
Unit-test the deterministic code around the model, run a golden set of real inputs through the feature with code-based graders where possible, and gate CI on the pass rate. Add LLM-as-judge only for qualities code cannot check, and monitor production to grow the set.
What is a golden set in LLM testing?
A fixed, versioned list of inputs with properties a correct output must have or avoid. It works as the feature's regression suite. Build it from real inputs and edge cases, and add every production failure to it.
Is LLM-as-judge reliable?
Partly. Zheng et al. found strong judges agree with human preferences over 80% of the time, but judges show position, verbosity and self-enhancement bias. Use pass/fail or pairwise questions, a different model, and check the judge against human labels.
How do you stop flaky LLM tests from failing CI?
Gate on a pass rate across many cases, not on single outputs. Set the threshold slightly below a measured baseline, pin the model version, lower the temperature, and report failed case IDs so a person can review them.
How many test cases does an LLM eval need?
Start with 50 to 100 real cases that run in minutes, including edge cases. Anthropic's guidance favours more cases with automated grading over fewer hand-graded ones. Grow the set from production logs.
Should evals run on every pull request?
Run them on every change that can alter model behaviour: prompts, model settings, retrieval and the AI code. Filtering CI by file path keeps ordinary pull requests fast and limits API cost.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.