Back to BlogAI Features & LLMs

What an AI Tutor Really Costs per Student per Month

Rupak Amin

Founder & Lead Engineer, RAITHub

13 min read

A typical AI tutor student costs about $1.53 a month in model fees, assuming 12 sessions of 15 turns, prompt caching and a 70/20/10 split across Claude Haiku 4.5, Sonnet 5.5 and Opus 5.5 at list prices. A light student costs about $0.51 and a heavy one about $5.11. Embeddings are a rounding error; output tokens, uncached context and the premium tier drive the bill.

If you would rather have it built for you, see how RAITHub would build this below.

The numbers in this post are illustrative, built from public list prices checked on 2 October 2026 and stated assumptions you can change. They are not PadhAI's costs: RAITHub built PadhAI, an AI tutoring platform with a 70/20/10 model router, and does not publish its usage or cost per student. The wider platform picture is in the EdTech software development guide.

How many tokens does one AI tutoring session use?

It depends on four things: the fixed instructions sent with every turn, the retrieved study material, the conversation so far, and the answer. A Socratic tutor, one that asks guiding questions instead of handing over answers, has short replies and many turns, so the input side dominates.

Part of each turnAssumed tokensCacheable?Notes
System prompt, tutoring rules, examples2,000YesSame for every student, so the cache stays warm under traffic
Retrieved material (RAG)1,500RarelyThree to five passages from the syllabus or textbook
History and the student's question2,000PartlyAn average; it grows during a session, so cap or summarise it
Tutor's reply400NoOutput, including any thinking tokens the model uses

That is 2,000 cached and 3,500 fresh input tokens, plus 400 output tokens, per turn. Measure your own: Anthropic's pricing page puts one token at roughly 4 characters or 0.75 words in English, and says "the exact count varies by language and content type", so Bangla, Hindi or Urdu text can count differently. It also notes that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text" (Anthropic pricing).

Which AI model is lowest-cost for a tutor, per turn?

Here is the per-turn cost of the turn above on current models from three providers, at standard list prices with the 2,000-token prefix read from cache.

Provider and model IDInput / cached input / output, per million tokensCost per turn
Anthropic, claude-haiku-4-5$1 / $0.10 / $5$0.0057
Anthropic, claude-sonnet-5-5$2 / $0.20 / $10$0.0114
Anthropic, claude-opus-5-5$4 / $0.20 / $20$0.0224
OpenAI, gpt-6-luna$0.10 / $0.01 / $0.50$0.00057
OpenAI, gpt-6.1-sol$2 / $0.10 / $10$0.0112
OpenAI, gpt-6-astra$10 / $1 / $50$0.0570
Google, gemini-3.1-flash-lite$0.25 / $0.025 / $1.50$0.0015
Google, gemini-3.8-flash$0.75 / $0.075 / $3.75 (through 31 December 2026)$0.0043
Google, gemini-3.1-pro-preview$2 / $0.20 / $12 (prompts up to 200k tokens)$0.0122

Sources: Anthropic pricing and model IDs, OpenAI API pricing and models, and Gemini API pricing, all checked on 2 October 2026. Three cautions. Google lists Gemini 3.8 Flash at $1.50 input and $7.50 output from 1 January 2027, which doubles that row. Anthropic's models page commits to Claude Haiku 4.5 only "not sooner than October 15, 2026" for retirement, so plan for the light tier to change model soon. And token counts differ between providers, so the same turn will not be exactly 5,900 tokens everywhere.

Price per turn is not the whole decision. A tutor that gets a maths step wrong costs you a student's trust, which is why the light tier should handle only the turns your evaluation set shows it handles well.

How much does prompt caching save an AI tutor?

A lot on the fixed part of the prompt. On Anthropic, a cache read costs 0.1 times the base input price on most models and 0.05 times on Claude Opus 5.5, while a 5-minute cache write costs 1.25 times. The pricing page says caching "pays off after one cache read for the 5-minute duration" (Anthropic pricing). OpenAI and Google list cached-input prices on their pricing pages too.

On Claude Sonnet 5.5, the 2,000-token prefix costs $0.004 per turn uncached and $0.0004 cached. Over 180 turns a month for a typical student, that is $0.72 against $0.07 on that part of the prompt alone. Two habits make it work: put everything that never changes at the start of the prompt, and keep per-student details after it.

What does a 70/20/10 model router do to cost per student?

It sends about 70% of turns to a lightweight model, 20% to a mid-tier model and 10% to a premium one, by predicting how hard each turn is. With the Anthropic per-turn costs above, the blended cost is 0.7 × $0.0057 + 0.2 × $0.0114 + 0.1 × $0.0224, or about $0.0085 a turn. That is about 25% below sending everything to Sonnet 5.5 and about 62% below sending everything to Opus 5.5. How to design and calibrate the router is in LLM model routing.

PadhAI's router classifies query complexity and sends each query to a lightweight, mid-tier or premium model. In a tutor, the premium share is where accuracy matters most: multi-step problems, misconceptions and marking explanations. PadhAI also checks mathematical results with code rather than trusting any tier, which is far cheaper than a second model call.

What does an AI tutor cost per student per month: light, typical and heavy?

Using the turn above, 15 turns per session, prompt caching and the 70/20/10 Anthropic mix:

StudentSessions a monthTurns a monthRouted 70/20/10All on Sonnet 5.5All on Opus 5.5
Light460$0.51$0.68$1.34
Typical12180$1.53$2.05$4.03
Heavy40600$5.11$6.84$13.44

The heavy row is the one to plan for. If your plan is a flat monthly fee, a small group of heavy users can take most of the AI budget. Common controls are a fair-use cap on premium-tier turns, summarising long histories, and moving non-urgent work such as weekly progress reports to batch processing, which Anthropic and OpenAI both discount by about 50%. For comparison, Khan Academy lists Khanmigo for learners and parents at $4 a month (Khanmigo pricing), a useful reference for what families already pay for an AI tutor.

Here is the same calculation as code, so you can change the assumptions.

type Price = { inPerM: number; cachedPerM: number; outPerM: number }
type Turn = { cachedIn: number; freshIn: number; out: number }

// List prices checked 2 Oct 2026. Keep them in config with a "checked on" date.
const PRICES: Record<string, Price> = {
  'claude-haiku-4-5': { inPerM: 1, cachedPerM: 0.1, outPerM: 5 },
  'claude-sonnet-5-5': { inPerM: 2, cachedPerM: 0.2, outPerM: 10 },
  'claude-opus-5-5': { inPerM: 4, cachedPerM: 0.2, outPerM: 20 },
}
const MIX: Record<string, number> = {
  'claude-haiku-4-5': 0.7,
  'claude-sonnet-5-5': 0.2,
  'claude-opus-5-5': 0.1,
}

function turnCost(p: Price, t: Turn): number {
  return (t.cachedIn * p.cachedPerM + t.freshIn * p.inPerM + t.out * p.outPerM) / 1_000_000
}

export function monthlyCostPerStudent(sessions: number, turnsPerSession: number, t: Turn): number {
  const blended = Object.entries(MIX).reduce((sum, [model, share]) => sum + share * turnCost(PRICES[model], t), 0)
  return blended * turnsPerSession * sessions
}

// Typical student: about 1.53 (USD)
monthlyCostPerStudent(12, 15, { cachedIn: 2000, freshIn: 3500, out: 400 })

How much do RAG and embeddings add to AI tutor cost?

Very little in fees; the real cost of RAG, retrieval-augmented generation, is the retrieved text you send to the model on every turn. Embedding a course library is a one-off: 10,000 pages of 500 tokens is 5 million tokens, about $0.10 with OpenAI's text-embedding-3-small at $0.02 per million or $0.65 with text-embedding-3-large at $0.13 (OpenAI pricing), or about $1 with Google's gemini-embedding-2 at $0.20 per million (Gemini pricing). Embedding each student question is fractions of a cent a month.

The retrieved passages are different. In the turn above, 1,500 RAG tokens are about 40% of the fresh input. Retrieving three good passages instead of eight mediocre ones is a cost decision as well as a quality one. For when RAG is the right tool at all, see RAG vs fine-tuning.

What does hosting add per student?

Usually less than the model bill once you have a few hundred active students. Two example list prices: Vercel Pro is $20 a month per developer seat with a $20 usage credit (Vercel pricing), and Neon's Launch plan charges $0.106 per compute-unit hour and $0.35 per GB-month of storage, with no monthly minimum (Neon pricing). One compute unit running all month is about 730 hours, or roughly $77; a database that scales down when idle costs less. A vector index can live in the same Postgres database, so retrieval does not need a separate paid service at this size.

Monthly item, illustrative1,000 typical studentsPer student
Model fees, routed 70/20/10About $1,530$1.53
App hosting and databaseAbout $100 to $300$0.10 to $0.30
Embeddings for questions and new materialUnder $5Under $0.01
TotalAbout $1,630 to $1,835About $1.63 to $1.84

Messaging channels, monitoring, email and payment fees come on top. If you sell in Bangladesh, collecting the subscription through local wallets is covered in integrating bKash, Nagad and SSLCommerz.

Buy, build or hire: should you build your own AI tutor?

OptionExample and cited priceChoose this whenWatch out for
Off-the-shelf AI tutorKhanmigo: $4 a month for learners and parents, free for teachers, district pricing on request (Khanmigo pricing)You need tutoring for your students, not a tutoring product of your ownYou cannot bring your own syllabus, brand or pricing
Template on a model APIA chat template wired to one model, paying list token prices such as gpt-6-luna at $0.10 input and $0.50 output per million (OpenAI pricing)A pilot with one subject and a few hundred studentsNo routing, retrieval quality checks or usage caps; cost and accuracy drift as you grow
Custom buildYour own tutor with routing, RAG on your material, caching and per-student cost trackingThe tutor is your product, you sell to many students, and margin per student mattersWeeks of engineering, plus ongoing evaluation as models and prices change

Doing the cost work yourself on an existing tutor, adding caching, a rules-first router and per-student cost logging, takes an experienced backend developer roughly 1 to 3 weeks. The main risk is quality: a router that sends hard maths to the light tier saves money while students quietly get wrong answers. Calibrate against an evaluation set first, as in how to test LLM features. Our guide to reducing LLM API cost lists the other levers.

Why RAITHub for this

  • Built an AI tutor end to end. PadhAI: 11 services (9 Node/TypeScript, 2 Python), a 70/20/10 LLM router, a Socratic tutor with math verification and RAG, PWA, WhatsApp and Telegram channels, and 9 payment gateways. The design is in how to build an AI tutor.
  • Cost measured, not guessed. Per-turn cost logging by model and feature, and routing changes judged against an evaluation set in CI.
  • Cost discipline elsewhere too. This site's own Neon database cost fix and rate limiting without Redis came from the same habit.
  • Your accounts, your data. We sign NDAs and DPAs and work inside your controls; production and student data stay in your own cloud account, and development uses synthetic data.

When you don't need us

  • You want tutoring for your own learners. An off-the-shelf tutor is faster and costs less than any build.
  • Your pilot is small. Under a few hundred students, one mid-tier model with prompt caching is enough; add routing when the bill justifies it.
  • You only need a cost estimate. The calculator above, with your own token logs, will answer it.

How RAITHub would build this

  • Scope: a tutor service with a cache-friendly prompt layout; a rules-first model router calibrated on your evaluation set; RAG over your syllabus with passage limits; per-student and per-feature cost logging with fair-use caps; a web app or PWA for students (web only, no native mobile apps).
  • Timeline: a new AI tutor MVP fits the 4 to 6 week fixed-scope range; adding routing, RAG and cost controls to an existing platform is backend and API work, 6 to 12 weeks.
  • You receive: automated tests and CI, including the evaluation set as a regression test; handover docs and runbooks for price changes and model swaps; and full IP under NDA.
  • Next step: a free 15-minute technical audit, then a written fixed quote.

See the EdTech industry page and the SaaS development service, or size a build with the MVP cost estimator. To model your own numbers, book the free 15-minute technical audit and bring a month of token logs if you have them.

Last reviewed: 2 October 2026. Anthropic, OpenAI, Google, Vercel, Neon and Khanmigo pricing checked on 2 October 2026.

Frequently asked questions

How much does an AI tutor cost per student per month?

On our worked example, about $0.51 for a light student, $1.53 for a typical one and $5.11 for a heavy one in model fees, with caching and a 70/20/10 router on Anthropic list prices. Hosting adds roughly $0.10 to $0.30 per student at 1,000 students.

How many tokens does an AI tutoring session use?

In our example, about 5,900 tokens per turn: 2,000 cached instructions, 1,500 retrieved material, 2,000 of history and question, and a 400-token reply. A 15-turn session is about 88,500 tokens.

Which model is lowest-cost for an AI tutor?

On list prices checked on 2 October 2026, small models such as gpt-6-luna and gemini-3.1-flash-lite cost a fraction of a cent per turn. Route only the turns they handle well to them, and send hard problems to a stronger tier.

Does prompt caching really cut AI tutor cost?

Yes, on the fixed part of the prompt. On Claude Sonnet 5.5, a 2,000-token cached prefix costs $0.0004 per turn instead of $0.004. Put fixed instructions first so they are cacheable.

How much do embeddings cost for a tutor's RAG?

Very little. Embedding 5 million tokens of course material costs about $0.10 on text-embedding-3-small. The bigger RAG cost is the retrieved text sent to the model on every turn.

What is a 70/20/10 LLM router?

A step that sends about 70% of turns to a lightweight model, 20% to a mid-tier model and 10% to a premium one by predicting difficulty. PadhAI, built by RAITHub, uses this split as its design target.

ai tutor cost per studentai tutor costllm cost per userprompt cachingmodel routingedtech ai

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.