Founder & Lead Engineer, RAITHub
Use classic RAG, one retrieval and one model call, for questions a single search can answer, which is most support and documentation traffic. Use agentic RAG, where a model plans, searches several times and checks its evidence, only for multi-part questions, comparisons and questions that need live tools. In the worked example below, the agentic path costs about five times more per question and takes several sequential calls.
This guide is for teams that already have, or are planning, a retrieval-augmented chatbot or assistant and are being told it should be "agentic". It explains what that word changes in the architecture, what it costs, and how to decide per question rather than per product. If you have not built the basic pipeline yet, start with how to build a RAG chatbot on your own product data.
What is the difference between RAG and agentic RAG?
Classic RAG (retrieval-augmented generation) is a fixed pipeline: embed the question, fetch the closest passages, and ask the model to answer from them. Agentic RAG hands control of that pipeline to the model, which decides what to search for, whether the results are good enough, and when to stop.
Anthropic's engineering guide Building effective agents draws the same line in general terms. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths"; agents are "systems where LLMs dynamically direct their own processes and tool usage". Classic RAG is a workflow. Agentic RAG is an agent whose main tool is search.
The 2025 survey Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG (Singh et al.) describes agentic RAG as embedding autonomous agents into the RAG pipeline, and names four design patterns it relies on: reflection, planning, tool use and multi-agent collaboration.
| Aspect | Classic RAG | Agentic RAG |
|---|---|---|
| Who decides what to retrieve | Your code: one search on the user's question | The model: it writes its own queries, possibly several |
| Model calls per question | Usually one | Several, often sequential |
| Sources | One index, such as a pgvector table | Several indexes, SQL, APIs, web search |
| Self-checking | None, unless you add a separate verifier | The model judges whether the evidence answers the question |
| Latency | Roughly one model call plus a search | The sum of every step in the loop |
| Cost predictability | High: tokens per question barely vary | Low: depends on how many steps the model takes |
| Failure mode | Wrong or missing passage, confident answer | Loops, drift from the question, runaway cost |
| Testing | Retrieval hit rate plus answer grading | All of that, plus step traces and stop behaviour |
What does agentic RAG cost compared with plain RAG?
More tokens and more waiting, because each step re-sends the growing context. Anthropic's guide says it plainly: "Agentic systems often trade latency and cost for better task performance." The question is whether your users' questions need that trade.
An illustrative comparison on gpt-5-mini, priced at $0.25 per million input tokens and $2.00 per million output tokens on the OpenAI pricing page (checked 29 September 2026). The token counts are assumptions to show the arithmetic, not benchmarks.
| Path (illustrative) | Model calls | Input tokens | Output tokens | Cost per question | Per 100,000 questions |
|---|---|---|---|---|---|
| Classic RAG: instructions plus 5 passages | 1 | 3,000 | 400 | $0.00155 | $155 |
| Agentic RAG: plan, 3 search rounds, check, answer | 4 | 20,000 (2k + 4k + 6k + 8k) | 1,200 | $0.0074 | $740 |
The agentic path is about 4.8 times the cost here, and the four calls run one after another, so the user waits for all of them. If the agent also uses a hosted search tool, add that: the same OpenAI page lists web search at $10.00 per 1,000 calls and file search at $2.50 per 1,000 calls, on top of tokens. Three web searches per question would add $3,000 per 100,000 questions, far more than the tokens.
For a wider view of build and running cost for agents, see what it costs to build an AI agent.
Which questions actually need agentic RAG?
Questions where one search on the user's words cannot find everything the answer needs. Anthropic's guide describes agents as suited to "open-ended problems where it's difficult or impossible to predict the required number of steps". In a product assistant, that usually means five kinds of question.
- Multi-part questions. "Does the Pro plan include SSO, and how does that compare with what we pay now?" needs a plan document and the customer's billing record.
- Comparisons across documents. "What changed in the refund policy between 2025 and 2026?" needs two versions retrieved on purpose.
- Questions that need live data. Order status, stock or account limits belong in a database or API call, not in an embedding index.
- Vague first questions. An agent can rewrite "it's not working" into several targeted searches, or ask a clarifying question.
- Research-style tasks for internal users, such as "summarise every incident report that mentions the payments service".
Most traffic is none of these. "How do I reset my password?" or "What file types can I upload?" is one search and one answer. Sending those through an agent pays five times the cost for the same reply.
Is there a middle ground between RAG and a full agent?
Yes, and it is where most products should start: fixed workflow steps that fix classic RAG's known weaknesses without handing the loop to the model.
- Query rewriting. One small model call turns the user's message, plus chat history, into a clean search query. Fixed cost, one extra call.
- Hybrid search. Keyword plus vector search in the same query catches exact terms such as error codes and plan names that embeddings blur.
- Re-ranking. Retrieve 20 passages, keep the 5 most relevant. Better context, no extra reasoning loop.
- A fixed second retrieval. If the top passage scores below a threshold, run one more search with a rewritten query, then answer or say "I don't know".
- Structured lookups by code. Detect intents like "order status" and call the API directly, with no model deciding.
Each of these is a predictable, testable step. Only when they still fail on a measurable share of questions is an agent worth its cost. That is the order Anthropic's guide recommends too: "find the simplest solution possible, and only increasing complexity when needed."
How do you route between classic and agentic RAG?
Classify each question first, send the simple ones down the fixed pipeline, and give the agentic path a hard step limit. The sketch below shows the shape; classify, search, judge and answer are your own functions.
type Passage = { id: string; text: string; score: number }
type Route = 'simple' | 'agentic'
declare function classify(question: string): Promise<Route>
declare function search(query: string, k: number): Promise<Passage[]>
declare function judge(question: string, evidence: Passage[]): Promise<{ enough: boolean; nextQuery?: string }>
declare function answer(question: string, evidence: Passage[]): Promise<string>
const MAX_ROUNDS = 3 // hard cap on agentic search rounds
export async function respond(question: string): Promise<string> {
if ((await classify(question)) === 'simple') {
return answer(question, await search(question, 5)) // classic RAG: one search, one answer
}
const evidence: Passage[] = []
let query = question
for (let round = 0; round < MAX_ROUNDS; round++) {
for (const p of await search(query, 5)) {
if (!evidence.some((e) => e.id === p.id)) evidence.push(p) // no duplicate passages
}
const verdict = await judge(question, evidence)
if (verdict.enough || !verdict.nextQuery) break
query = verdict.nextQuery
}
return answer(question, evidence) // the answer step must still refuse if evidence is thin
}
Three details matter. The classifier can be a small, cheap model or plain rules, because it only picks a lane. The step cap is enforced by code, not requested in the prompt; Anthropic's guide notes that agents commonly include "stopping conditions (such as a maximum number of iterations) to maintain control". And the final answer step keeps the same "answer only from the sources" rule as the simple path, so more searching never becomes licence to guess. Spend limits and timeouts for the agentic lane are a subject of their own.
How do you know whether agentic RAG is helping?
Run the same evaluation set through both paths and compare accuracy, cost and latency per question type. If the agent does not beat the simple path by a margin you can name on the question types you route to it, it is not worth running.
| Measure | What it tells you |
|---|---|
| Answer accuracy per question type | Whether the agent earns its cost on the questions routed to it |
| Retrieval recall | Whether the right passage was found at all, by either path |
| Steps per question | Whether the agent is looping, and how often it hits the cap |
| Cost and latency per question | The real multiplier in your traffic, not the illustrative one above |
| Misroutes | Simple questions sent down the agentic lane, and the reverse |
The method for building that evaluation set is in how to test LLM features. If the gap is about knowledge rather than retrieval, the choice may be a different one entirely; see RAG vs fine-tuning for startups.
What does RAITHub use in practice?
PadhAI, the AI tutoring platform RAITHub built, uses RAG so its Socratic tutor explains from approved source material, and a 70/20/10 model router that sends each query to a lightweight, mid-tier or premium model by complexity. The same principle applies here: decide per request how much machinery it needs, rather than paying the heaviest path for every message. RAITHub does not publish PadhAI's traffic or cost figures, and the routing code above is illustrative, not PadhAI's.
Why RAITHub for a RAG or agentic RAG build
- Built retrieval into a real product. PadhAI's tutor combines RAG with math verification across 11 services, 9 Node/TypeScript and 2 Python.
- Cost as a design input. A router that sends each request to the least expensive path that handles it, as in PadhAI's 70/20/10 design.
- Tested, not demoed. The same habit behind 750+ tests on TheSkinProof, the founder's own marketplace venture, and 530+ on Sundor Skin.
- Fixed scope, your IP. A free 15-minute technical audit, then a fixed written quote. You own the code and IP; an NDA is standard.
When you don't need us
- Your assistant answers from a small, stable set of docs. A hosted file search tool or a chatbot widget may be enough, with no agent at all.
- Your team already runs an evaluation set. Then the routing decision above is a measured experiment you can run yourselves.
- You want an autonomous agent with no limits. That is not a design RAITHub recommends: every agent loop needs step, token and spend caps enforced in code.
AI features inside a product are built under the SaaS development service. To decide which path your questions need, book the free 15-minute technical audit and bring 20 real user questions, including the ones your current assistant gets wrong.
Last reviewed: 29 September 2026. Prices and sources checked on 29 September 2026.
Frequently asked questions
What is agentic RAG in simple terms?
Retrieval-augmented generation where the model, not fixed code, decides what to search for, whether the results are good enough, and when to stop. It can search several times and use other tools before answering.
Is agentic RAG better than normal RAG?
Only for questions a single search cannot answer, such as multi-part questions, comparisons and questions needing live data. For simple lookups it gives the same answer at several times the cost and latency.
How much more does agentic RAG cost?
It depends on the number of steps. In an illustrative four-call loop on gpt-5-mini, it cost about 4.8 times a single-call RAG answer, before any hosted search tool fees, which can dominate.
Can I use both RAG and agentic RAG in one product?
Yes. Classify each question, send simple ones through classic RAG, and send only complex ones to a bounded agentic loop with a hard step limit.
What should I try before building agentic RAG?
Query rewriting, hybrid keyword and vector search, re-ranking, and one fixed fallback search. They fix most retrieval misses at a predictable cost.
How do I stop an agentic RAG loop from running forever?
Enforce a maximum number of rounds in code, deduplicate retrieved passages, and add token and spend limits and a timeout per request. Never rely on the prompt alone to stop the loop.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.