Founder & Lead Engineer, RAITHub
RAITHub ships and tests production software. See QA as a Service or talk to us.
AI chatbot and LLM feature testing as a service is QA for the parts of your product a language model drives: a provider builds a golden set of real inputs, scores outputs, tests guardrails and jailbreaks, and gates releases on a pass rate. You buy it as an audit, a plan or a team. It suits teams shipping a chatbot with no eval harness.
If you would rather have it tested for you, see how RAITHub would test this below, or go to the AI-built app testing service and the QA as a service overview. LLM means large language model; an eval is a repeatable test that runs the feature over a set of inputs and scores the outputs.
Why can't you test a chatbot with a demo?
Because a model's output varies from run to run and changes when the prompt or model changes, so a demo that worked once proves nothing about the next release. The same prompt can give a different answer tomorrow, and a prompt tweak that fixes one case can quietly break ten others. Testing has to measure a pass rate across many inputs, not a single happy path. OpenAI's own guidance calls evals one of the only ways to improve an LLM application and a way to test it despite this variability (OpenAI evaluation best practices).
The stakes are real. In Moffatt v. Air Canada (2024 BCCRT 149), a tribunal held an airline responsible for a refund its support chatbot had wrongly promised. A chatbot speaks for your company, and an untested one speaks without supervision. For the engineering method in full, see how to test LLM features; this post is the commercial service wrapped around it.
What does an LLM testing service cover?
| Area | What it answers | How it is tested |
|---|---|---|
| Deterministic code | Does prompt assembly, parsing and tool calling work for odd inputs? | Ordinary unit tests, no model call |
| Answer quality | Does the model answer correctly across real inputs and edge cases? | A golden-set eval with a pass-rate gate |
| Faithfulness | Does it stick to your data instead of inventing facts (hallucinating)? | Grounding checks against the retrieved source |
| Guardrails | Does it refuse out-of-scope, unsafe or off-brand requests? | Adversarial and jailbreak fixtures |
| Cost and limits | Can a prompt loop or oversized input run up unbounded cost? | Limit and timeout tests |
Most defects live in the deterministic parts, so those are tested first, with cheap, fast unit tests that never call a model. The model's answer quality is the part that needs evals. For the guardrail side, OWASP's LLM10:2025 Unbounded Consumption warns about resource and cost exhaustion, which is why limits and timeouts are tested, not assumed.
How does a golden-set eval work in code?
A golden set is a fixed, versioned list of inputs with properties a correct output must have or avoid. Run the feature over it, grade with code-based checks wherever possible, and fail the build if the pass rate drops below a baseline. Use the least expensive grader that is reliable; Anthropic's guide to building evaluations orders them exact match, string checks, code-graded checks, then model-based grading. A minimal assertion with Vitest:
Free checklist
AI-Built App Launch Readiness Checklist
25 checks before you let real users in. Enter your email and we’ll reveal it below (and send you a copy).
One email, the checklist, no spam. By submitting you agree we can email you this checklist and reply to your enquiry.
// evals/support-answer.eval.test.ts
import { readFileSync } from 'node:fs'
import { describe, expect, it } from 'vitest'
import { answerSupportQuestion } from '../src/ai/support-answer'
interface GoldenCase {
id: string
input: string
mustInclude: string[]
mustNotInclude: string[]
}
const cases: GoldenCase[] = JSON.parse(
readFileSync('evals/golden/support-answer.json', 'utf8'),
)
// Set from a measured baseline, not a guess. Raise it as the feature improves.
const MIN_PASS_RATE = 0.9
function passes(output: string, c: GoldenCase): boolean {
const text = output.toLowerCase()
return (
c.mustInclude.every((s) => text.includes(s.toLowerCase())) &&
c.mustNotInclude.every((s) => !text.includes(s.toLowerCase()))
)
}
describe('support answer golden set', () => {
it('meets the pass-rate gate', async () => {
const failed: string[] = []
for (const c of cases) {
const output = await answerSupportQuestion(c.input)
if (!passes(output, c)) failed.push(c.id)
}
const passRate = (cases.length - failed.length) / cases.length
expect(passRate).toBeGreaterThanOrEqual(MIN_PASS_RATE)
}, 300_000)
})
Where the model's output can be verified by computation, a lookup or a query, verify it in code rather than asking another model. PadhAI, the AI tutoring platform RAITHub built, runs a Socratic tutor whose mathematical claims are checked rather than trusted, and a 70/20/10 model router. When only a judgement call like tone or faithfulness is left, an LLM-as-judge can score it, but it has measured biases: strong judges agree with human preferences over 80% of the time, not always, and show position, verbosity and self-enhancement bias (Judging LLM-as-a-Judge, Zheng et al., 2023). Calibrate the judge against human labels first and pin its model version.
Buy, build or hire?
| Route | What it costs on the market | Choose this when | Watch out for |
|---|---|---|---|
| An eval tool or platform | Open-source eval runners are free; hosted platforms are quoted per scope | Your team will own the golden set, graders and thresholds | A tool runs evals; it does not decide what good looks like or read the failures |
| Freelancers | Upwork lists a median of $35 an hour for QA engineers (Upwork QA engineer rates) | A short eval build for one feature | LLM testing skill varies widely; the golden set leaves with them |
| An in-house AI QA hire | The US median wage for QA analysts and testers was $104,300 in May 2025, before benefits (US Bureau of Labor Statistics) | AI is core to the product and you can recruit a scarce skill | One hire rarely covers evals, guardrails and the deterministic code too |
| A managed LLM testing service | Quoted per scope: an audit, a plan or a dedicated team | You ship a chatbot or LLM feature and have no eval harness | Make sure the golden set and tests live in your repository |
How RAITHub would test this
- Scope: name the LLM features, the facts they must be faithful to, and the requests they must refuse.
- Golden set: 50 to 100 real inputs and edge cases with properties each output must have or avoid, built from your logs and versioned in your repository.
- Graders: code-based checks first, LLM-as-judge only for tone and faithfulness, calibrated against human labels.
- Guardrails: adversarial and jailbreak fixtures, plus limit and timeout tests against unbounded cost.
- CI gate: the eval set run on every change to prompts, model settings or AI code, failing the build below a measured baseline.
Timeline: a one-off audit is fixed in scope and dates before it starts; a plan or dedicated team runs month to month. You receive: the golden set, graders and CI configuration in your repository, bug reports in your tracker, a handover document, full IP and an NDA. Next step: a free 15-minute audit, then a written fixed quote. The QA proof RAITHub cites is counted test suites plus PadhAI's eval-first design; RAITHub does not publish PadhAI's eval scores, so treat method as the claim, not results. PropDesk runs 1,024 tests and this website runs 400+ in CI. There is no AI testing case study yet.
To start, tell RAITHub about your chatbot or LLM feature. For a single pre-launch pass, use the one-off audit form. For the chatbot grounding side, see fixing a chatbot that hallucinates about your product.
When don't you need RAITHub for this?
- Your team already keeps a golden set and CI eval gate. You may only need a review.
- The feature is low-stakes and always edited by a person before it ships. Unit tests and a small sampled review may be enough.
- You need a model red-team certification or a safety attestation. RAITHub's testing produces no certification.
- You want testers under your own management. RAITHub does not offer staff augmentation.
Frequently asked questions
What is LLM or AI chatbot testing as a service?
QA for LLM-driven features bought as a managed service: the provider builds a golden set of real inputs, grades outputs with code-based checks first, tests guardrails and jailbreaks, and gates releases on a pass rate, as a one-off audit, a monthly plan or a dedicated team.
How do you test a chatbot that gives different answers each time?
By measuring a pass rate across many inputs, not a single output. A versioned golden set runs on every change, graded by code-based checks where possible, with the build failing if the pass rate drops below a measured baseline.
Can you test for hallucinations and jailbreaks?
Yes. Faithfulness is checked by grounding the answer against the retrieved source, and guardrails are tested with adversarial and jailbreak fixtures plus limit and timeout tests against unbounded cost, following OWASP's LLM risk guidance.
Is LLM-as-judge reliable enough to grade a chatbot?
Partly. Strong judges agree with human preferences over 80% of the time but show position, verbosity and self-enhancement bias. Use it only for tone and faithfulness, prefer code-based graders, calibrate against human labels, and pin the judge model version.
Does the eval suite stay with us?
Yes. The golden set, graders and CI configuration live in your repository and the IP is assigned to you, so the suite keeps protecting you after the engagement ends.
How much does AI feature testing cost?
It depends on the number of features, the size of the golden set and the guardrail scope. Marketplace QA engineers typically charge $20 to $60 an hour. RAITHub publishes no rates and quotes a fixed price after a free audit.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.