Back to BlogAI Features & LLMs

Your AI Chatbot Hallucinates About Your Own Product: How to Fix It

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

A chatbot hallucinates about your product when it answers without the right source in front of it, or is allowed to fill gaps from general knowledge. Fix it: log what was retrieved for every answer, restrict the model to those sources with permission to say "I don't know", read prices and limits live from your database, and reject answers whose citations do not check out in code.

A hallucination is a confident answer that is not supported by fact: a plan tier you do not sell, a feature you retired, a refund window you never offered. On a product bot it is worse than a wrong answer on a general chatbot, because customers reasonably take it as your word. This guide is for teams whose assistant is live and saying things it should not. It assumes a retrieval-augmented generation (RAG) setup, where the bot fetches passages from your content for each question, and shows how to find and fix the failure.

Why does an AI chatbot make things up about your own product?

Because language models are built to produce a plausible answer, and nothing in a basic setup stops them when the real answer is missing. In Why Language Models Hallucinate (Kalai et al., 2025), OpenAI and Georgia Tech researchers argue that training and evaluation reward guessing over acknowledging uncertainty. Your bot inherits that habit unless you design it out.

In a product bot the guess usually has one of a few concrete causes, and each has a different fix:

SymptomLikely causeFix
Invents a feature or plan that sounds rightNothing relevant was retrieved, and the model filled the gapRefuse when retrieval is empty or weak; restrict to sources
Describes an old version of the productRetired or outdated docs still in the indexRe-index on change; delete retired sources
Wrong price, limit or stock levelLive facts copied into documents and embeddedRead them from the database at answer time
Mixes up two plans or two productsChunks without context, such as a table row with no plan nameCarry the title and section into every chunk; add metadata filters
Right passage retrieved, wrong answer givenThe prompt lets the model add general knowledgeAnswer only from sources, cite each claim, verify citations
Answers about a competitor or an unrelated topicNo scope limitDecline out-of-scope questions with a fixed reply

How do you find out why a specific answer was wrong?

Look at what the model saw. Store, for every answer, the question, the retrieved chunks with their scores, the final prompt, the model and the reply. Without that log you are guessing about a guess.

Then trace each bad answer through three questions:

  1. Was the right passage retrieved? If not, it is a retrieval problem: chunking, embeddings, filters or a missing document.
  2. Was the passage itself correct and current? If not, it is a content problem, and no model change will fix it.
  3. Was it retrieved and correct, but the answer still wrong? Then it is a generation problem: the prompt, the model or the missing checks below.

Sort twenty bad answers this way and the dominant cause is usually obvious. The build of the retrieval pipeline itself, chunking, pgvector and tenant filters, is outside this post; the wider pattern for shipping an AI feature safely is in how to add AI features to an existing SaaS.

How do you stop the chatbot answering beyond its sources?

Tell it to use only the provided sources, cite each claim, and say it does not know otherwise. Then enforce the empty case in code, so a question with no good match never reaches the model at all.

Anthropic's guide to reducing hallucinations lists these as the basic strategies: explicitly allow the model to say "I don't know", restrict it to the provided documents rather than its general knowledge, and have it cite a supporting quote for each claim, retracting any claim it cannot support. It also cautions that these techniques reduce hallucinations but do not eliminate them, which is why the checks below exist.

Make "I don't know" a good outcome in the product, not a failure: offer the closest help articles and a route to a person. A bot that refuses one question in ten and is right on the rest is worth more than one that answers everything and is wrong on one in ten.

Should prices and plan limits come from documents or the database?

From the database, every time. Anything that changes by the day, such as prices, plan limits, stock, delivery dates or a customer's own account state, should be fetched live through a tool call rather than embedded as text that goes stale.

import type { Pool } from 'pg'

// Exposed to the model as a tool. The answer quotes this row, not a document.
export async function getPlanFacts(db: Pool, planSlug: string) {
  const { rows } = await db.query(
    'SELECT name, monthly_price_usd, seat_limit, updated_at FROM plans WHERE slug = $1 AND is_active',
    [planSlug],
  )
  return rows[0] ?? null // null: the plan does not exist, and the bot must say so
}

The rule is simple: documents explain, the database states facts. A help article can say how seats work; only the plans table says how many seats the Growth plan has today. If a price appears in both, delete it from the document, or the index will eventually disagree with the checkout.

How do you check an answer before the customer sees it?

Verify it in code. A cheap deterministic check catches the worst failures: answers with no citation, citations to sources that were never retrieved, and numbers that appear in none of the cited sources.

type Source = { n: number; content: string }

// Flags an answer that cites nothing, cites a source that was not retrieved,
// or states a number that none of its cited sources contains.
export function groundingProblems(answer: string, sources: Source[]): string[] {
  const problems: string[] = []
  const byN = new Map(sources.map((s) => [s.n, s.content] as const))
  const cited = [...answer.matchAll(/\[(\d+)\]/g)].map((m) => Number(m[1]))
  if (cited.length === 0) problems.push('no citations')
  for (const n of cited) if (!byN.has(n)) problems.push('cites unknown source ' + n)

  const citedText = cited.map((n) => byN.get(n) ?? '').join(' ')
  const body = answer.replace(/\[\d+\]/g, '')
  for (const num of body.match(/\d+(?:[.,]\d+)*/g) ?? []) {
    if (!citedText.includes(num)) problems.push('number not in cited sources: ' + num)
  }
  return problems
}

If the list is not empty, regenerate once with the problems listed in the prompt, and if it still fails, show the fallback reply with links to the sources. Numbers are the check that pays off most, because wrong prices, limits and time windows are the hallucinations that cost money. Provider features help here too: Anthropic's Citations feature returns the exact cited passage with each claim, and its docs say citations are guaranteed to contain valid pointers to the provided documents.

This is the same principle as math verification in PadhAI, the AI tutoring platform RAITHub built: its Socratic tutor checks a mathematical result in code rather than accepting the model's word, and uses RAG so explanations draw on source material. Checks in code, not trust in the model, is what makes an AI answer dependable. How that works in a tutor is in how to build an AI tutor.

How do stale and conflicting docs cause hallucinations?

If the index holds two versions of the truth, retrieval will sometimes pick the wrong one, and the model will cite it faithfully. The answer is correct to its source and wrong for your customer.

  • Re-index on publish, edit and delete, not on a weekly schedule. Deletions matter most: a retired feature's doc is a standing invitation to describe it.
  • Tag chunks with plan, region and product version, and filter retrieval by the asking customer's context.
  • Keep one source of truth per fact. Old blog posts and release notes often contradict current docs; exclude them or label them as history.
  • Show the source date alongside each cited article, so customers and support staff can spot an old one.

What does a wrong chatbot answer cost a business?

Potentially what the bot promised. In Moffatt v. Air Canada, 2024 BCCRT 149, decided on 14 February 2024, British Columbia's Civil Resolution Tribunal held the airline responsible for its website chatbot's wrong description of its bereavement fare policy. The tribunal rejected the argument that the chatbot was a separate entity responsible for its own actions, and ordered the airline to pay about CA$812 in damages, interest and fees.

The amount was small; the principle is not. Treat bot answers as statements your business makes. This is general information, not legal advice; confirm what applies to your product and markets with your adviser.

How do you test a chatbot for hallucinations before and after launch?

With trap questions, not just happy-path ones. A golden set that only asks questions your docs answer will never show you a hallucination.

  • Answerable questions with the expected fact and source.
  • Unanswerable questions your docs deliberately do not cover. The pass is "I don't know".
  • Stale traps: questions about retired features and old prices. The pass is the current answer or a refusal.
  • Near-miss questions that mix two plans or products.
  • Out-of-scope and injection attempts: competitors, unrelated topics, "ignore your instructions".

Run the set on every change to prompts, models, chunking or content, and gate releases on it. Track the grounding-check failure rate in production as a live metric. RAITHub's approach to CI gates is in how RAITHub tests software.

Why RAITHub for fixing a hallucinating chatbot

  • Verification as a habit. PadhAI, which RAITHub built, checks math in code and grounds its tutor with RAG; the same approach applies to your product's facts.
  • Diagnosis before rewrite. RAITHub traces bad answers through the logs first, and fixes the cause, whether retrieval, content or generation.
  • Tested releases. The discipline behind 750+ tests on TheSkinProof, the founder's own marketplace venture, and 530+ on Sundor Skin, applied as a trap-question gate.
  • Fixed scope. A free 15-minute technical audit, then a fixed written quote. You own the code and the IP, and an NDA is standard.

When you don't need us

  • The bot is a vendor product. If your help desk vendor runs the bot, raise it with them and clean up the help centre it reads from.
  • The cause is content. If your trace shows the docs are wrong or stale, a content owner fixes that faster than an engineer.
  • You need a legal view on what your bot has already promised. That is a question for your adviser, not a developer.

For a live bot giving wrong answers, see how RAITHub handles a specific fix, or the QA and test automation service for the evaluation gate. To start, book the free 15-minute technical audit and bring five wrong answers with the logs behind them.

Last reviewed: 29 September 2026. Research and docs checked on 29 September 2026.

Frequently asked questions

Why does my AI chatbot give wrong answers about my product?

Usually because the right passage was not retrieved, the content itself is stale, live facts such as prices were embedded as text, or the prompt lets the model fill gaps from general knowledge. The retrieval log tells you which.

Can you stop an AI chatbot from hallucinating completely?

No. Grounding, citations and code checks reduce it sharply, and Anthropic's own guide says these techniques do not eliminate it. Design a safe fallback for the cases that slip through.

Should my chatbot say "I don't know"?

Yes. Explicit permission to admit uncertainty is one of the basic strategies in Anthropic's hallucination guide, and a refusal with links to help articles is safer than a confident invention.

How do I stop my chatbot quoting wrong prices?

Remove prices from the documents it retrieves, and have it fetch them from your database through a tool call at answer time. Then check that any number in the answer appears in its cited source.

Is a company responsible for what its chatbot says?

In Moffatt v. Air Canada (2024), a Canadian tribunal held the airline responsible for its chatbot's wrong policy information. This is general information; confirm your position with your adviser.

How do I test a chatbot for hallucinations?

Build a golden set that includes unanswerable questions, stale-content traps and near-miss questions, not only ones your docs answer, and run it on every prompt, model or content change.

AI chatbot hallucinationchatbot wrong answersRAG groundingcitationsLLM evaluationsupport chatbot

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.