Back to BlogAI Features & LLMs

AI Math Tutors That Don't Give Wrong Answers: Verification Loops

Rupak Amin

Founder & Lead Engineer, RAITHub

13 min read

An AI math tutor stops giving wrong answers when the language model never has the last word on correctness. The model reads the problem and the learner's work; code computes the truth with a computer algebra system or a sandboxed evaluator; the tutor confirms only what the checker agrees with. Models need that check: Apple researchers found one irrelevant clause cut accuracy by up to 65% (GSM-Symbolic).

If you would rather have the verification layer built for you, see how RAITHub would build this below.

This post goes one level deeper than our guide on how to build an AI tutor students can trust, which covers pedagogy, RAG and safety. Here the subject is narrower: the loop that sits between the model and the learner and decides whether "Correct!" is true.

Why do ChatGPT and other LLMs get maths wrong?

Because a language model predicts likely text; it does not run arithmetic or algebra. It is often right, because worked solutions are common in its training data, but "often" is not good enough when a learner takes the tutor's verdict as fact.

The research shows how brittle that pattern-matching is. In GSM-Symbolic (Mirzadeh et al., Apple), models showed noticeable variance when only the numbers in a grade-school word problem changed, performance fell as clauses were added, and adding a single clause that seemed relevant but did not affect the answer caused drops of up to 65% across the state-of-the-art models tested. The authors conclude that current models may be replicating reasoning patterns from training data rather than reasoning step by step.

For a tutor, three failure modes matter most:

  • Wrong final answer. An arithmetic slip, a sign error or a dropped root, stated with full confidence.
  • Wrong verdict on the learner. The model marks a correct answer as wrong, or an equivalent form (2(x + 3) versus 2x + 6) as different. This one damages trust fastest, because the learner was right.
  • Plausible but broken steps. The final answer is right but a middle step is invalid, and the learner copies the method.

How do you make an AI math tutor accurate?

Split the work. The model does language: reading the problem, turning it into a formal expression, explaining. Code does mathematics: computing the answer and checking every claim against it. This is the idea behind program-aided language models. In PAL (Gao et al.), the model writes a program and a Python interpreter solves it, and the approach beat a much larger chain-of-thought model on the GSM8K word-problem benchmark by 15 points of top-1 accuracy.

A verification loop for a tutor has four stages:

StageWhat happensWho does it
1. FormaliseThe problem becomes a machine-checkable form: an equation, an expression, or a short programThe model, or the question bank if problems are authored
2. Compute truthThe expected answer is computed, not generatedA computer algebra system (CAS) such as SymPy, or a sandboxed evaluator such as math.js
3. CheckThe learner's final answer and each step are compared with the truthDeterministic code, with three outcomes: correct, wrong, can't verify
4. RespondThe model writes the reply, given the verdict as a fixed fact it may not overrideThe model, then an output filter

Two design rules make this work. First, the model never sees an unverified "correct": the verdict is passed to it as data, and the prompt tells it to explain that verdict, not judge again. Second, if the problem comes from your own question bank, compute and store the answer when the question is authored, so stage 1 cannot go wrong at runtime. Model formalisation is for free-form questions learners type in.

What does a math answer checker look like in code?

For equations, the simplest reliable check is substitution: put the claimed solution back into the equation and see whether both sides agree. It needs no solver, and it works whether the claim came from the learner or from the model. Below is a minimal TypeScript version using math.js, locked down as the math.js security guide recommends:

import { create, all } from 'mathjs'

const math = create(all)
const evaluate = math.evaluate // keep a reference before disabling

// Disable the functions the math.js security guide flags as highest risk.
math.import(
  {
    import: () => { throw new Error('disabled') },
    createUnit: () => { throw new Error('disabled') },
    reviver: () => { throw new Error('disabled') },
  },
  { override: true },
)

export type Verdict = 'correct' | 'wrong' | 'unverifiable'

/** Is "claimed" a solution of the equation? Checked by substitution. */
export function checkSolution(equation: string, variable: string, claimed: number): Verdict {
  const sides = equation.split('=')
  if (sides.length !== 2 || !Number.isFinite(claimed)) return 'unverifiable'
  try {
    const scope = { [variable]: claimed }
    const diff = Number(evaluate(sides[0], scope)) - Number(evaluate(sides[1], scope))
    if (!Number.isFinite(diff)) return 'unverifiable'
    return Math.abs(diff) < 1e-9 ? 'correct' : 'wrong'
  } catch {
    return 'unverifiable' // unparseable input is never marked correct
  }
}

checkSolution('2*x + 3 = 11', 'x', 4)    // 'correct'
checkSolution('x^2 - 5*x + 6 = 0', 'x', 3) // 'correct'
checkSolution('x^2 - 5*x + 6 = 0', 'x', 4) // 'wrong'

Four production notes:

  • Run it in isolation. The math.js guide warns that evaluating arbitrary expressions carries risk and recommends a separate worker or child process with a timeout, so an expensive expression cannot freeze the server. On the Python side, the SymPy core reference warns that sympify uses eval and should not be used on unsanitised input. Never feed raw learner text to either in your main process.
  • Substitution proves a root, not completeness. If the learner says x = 3 for a quadratic with roots 2 and 3, each claimed root passes. Compare against the full solution set, computed by a CAS when the question is created, to catch a missing root.
  • Floating point needs a tolerance. The fixed 1e-9 above suits small integer problems; scale it for large magnitudes, and use exact rational arithmetic where fractions matter.
  • "Unverifiable" is a real outcome. The tutor says it cannot check that form and asks the learner to rewrite it, or walks through a check together. It never guesses.

For symbolic answers, such as checking that 2(x + 3) equals 2x + 6, use a CAS equivalence test instead; our AI tutor guide shows the SymPy version. This code is a minimal illustration of the technique, not PadhAI's code.

How do you check each step, not just the final answer?

Most learning happens at the step where things go wrong, so a tutor that only checks the final answer can only say "try again". Step-checking lets it say "your third line is where it changed".

For equation solving there is a cheap, deterministic method: the true solution should satisfy every line of a valid working. Substitute the known root into each of the learner's lines in turn; the first line it no longer satisfies is the first wrong step. For expression simplification, check that each line is equivalent to the one before with a CAS.

Know the edge cases. Squaring both sides can add solutions, and dividing by an expression that may be zero can lose them, so a line can be "satisfied" and still be an unsafe move. Flag those operations for the model to discuss rather than marking them simply right.

For steps that are not formal algebra, such as a geometry argument or a word-problem plan, there is no cheap deterministic check. Research points to step-level grading: in Let's Verify Step by Step (Lightman et al., OpenAI), a model trained with step-level feedback, called process supervision, solved 78% of a representative subset of the MATH test set and significantly outperformed feedback on final answers alone. That is a training-time technique for model builders; a product team can borrow the idea by having a second model grade each step and treating its judgement as advisory, never as a verified "correct".

How does a Socratic math tutor avoid giving the answer away?

By knowing the answer privately and enforcing the rule in code. A prompt that says "don't reveal the answer" is a request, and learners will try to talk the model out of it.

  • Keep the computed answer server-side. It lives in session state and is passed to the model only as context for hints, never in text the learner can extract.
  • Filter every reply. Before sending, code checks whether the reply contains the final answer, in any equivalent form, before the learner has reached it. If it does, regenerate with a stricter instruction or trim it.
  • Use a hint ladder. Nudge, then point to the idea, then a worked similar example with different numbers. Each rung is logged, so a teacher can see how much help was needed.
  • Confirm only verified work. When the learner reaches an answer, the checker's verdict decides what the tutor says.

That last point matters most: a Socratic tutor that praises a wrong answer has done more harm than a direct one, because the learner believed they had worked it out.

How long does it take to add math verification yourself?

For a developer comfortable with TypeScript or Python, plan on roughly 2–4 weeks to add final-answer and step checking for one topic band, such as arithmetic and linear equations, with an isolated evaluator and an evaluation set. Each new topic, such as quadratics, inequalities or calculus, adds parsing rules and edge cases.

The main risk is a checker that is confidently wrong: marking an equivalent answer as wrong, or running untrusted input in the main process. Both are found by testing against a fixed set of real learner answers, including messy ones, on every change. Our guide to testing LLM features covers how to build that set.

Should you buy a math tutor, build on a hosted tool, or hire a team?

OptionExample and costChoose this whenThe catch
Off-the-shelf tutorKhanmigo, $4 a month for learners and parents and free for teachers (Khanmigo pricing)You need a tutor for your learners, not a tutoring product of your ownNot your brand, your curriculum or your data model
Model plus hosted code executionOpenAI's Code Interpreter tool, from $0.03 per 20-minute 1 GB container session (OpenAI pricing)You are prototyping, or adding occasional calculation to a general assistantThe model decides when to run code; step-checking, answer withholding and "can't verify" are still yours to build
Custom verification layerYour own checker, evaluator sandbox and output filterCorrect verdicts are the product: a tutoring platform, exam prep or homework help at scaleBuild and maintenance time, and an evaluation set you must keep current

Prices checked 2 October 2026. Model choice is a separate decision; routing easy queries to lighter models is covered in LLM model routing for cost.

Why RAITHub for this

  • We have built one. PadhAI, the AI tutoring platform RAITHub built, has a Socratic tutor with math verification and RAG, and a 70/20/10 LLM router that sends each query to a lightweight, mid-tier or premium model.
  • The right language for each job. PadhAI runs as 11 services, 9 in Node/TypeScript and 2 in Python, so product logic and maths or AI work each sit where they fit.
  • One tutor on every channel. PadhAI serves a PWA, WhatsApp and Telegram, with 9 payment gateways, so a verification fix ships everywhere at once.
  • QA-first by habit. A verification layer is only as good as its tests; the same discipline produced 750+ tests on TheSkinProof, the founder's own venture.

RAITHub does not publish PadhAI accuracy figures or learning outcomes. The architecture is on the work page.

When you don't need us

  • Your learners just need a tutor. An existing product your school approves is faster and lower-cost than building one.
  • Your maths is all multiple choice. If every answer is a fixed option from an authored bank, you need answer keys, not a verification loop.
  • You need efficacy evidence from past work. RAITHub cannot show you published outcome data.
  • You need curriculum design. RAITHub builds the platform; subject experts write the questions and worked examples.

How RAITHub would build this

  • Scope: a verification service with final-answer and step checks for your first topic bands, running in an isolated worker with timeouts.
  • Scope: question-bank tooling that computes and stores the true answer and full solution set when a question is authored.
  • Scope: a Socratic reply pipeline with server-side answers, a hint ladder and an answer-leak filter.
  • Scope: an evaluation set of real learner answers, run in CI on every prompt, model or checker change.

Timeline: a tutoring MVP with verification is typically 4–6 weeks of fixed scope; adding a verification service to an existing platform is backend work in the 6–12 week range, depending on topics and channels.

You receive: automated tests and CI, handover docs and runbooks, and full IP under NDA. Learner data stays in your own cloud account, and development uses synthetic data.

Next step: a free 15-minute technical audit, then a written fixed quote. Book the audit and bring sample problems and the topics you need first. For the wider picture, see the EdTech software development guide, what RAITHub builds on the EdTech page, and the SaaS development service.

Last reviewed: 2 October 2026. Research and pricing checked on 2 October 2026.

Frequently asked questions

Why do AI tutors get maths wrong?

Language models predict likely text instead of calculating. Apple's GSM-Symbolic study found that changing numbers shifted results, and one irrelevant clause cut accuracy by up to 65% across the models tested.

Can an AI math tutor be 100% accurate?

Not on every question, but it can avoid confident errors. Compute answers with code, check learner answers deterministically, and say "can't verify" instead of guessing when an answer cannot be parsed or checked.

What is a computer algebra system?

Software that manipulates mathematical expressions exactly, such as solving equations or testing whether two expressions are equal. SymPy is a widely used open-source one for Python.

How do you check a learner's steps automatically?

For equations, substitute the known solution into each line; the first line it fails is the first wrong step. For simplification, test that each line is equivalent to the previous one with a CAS.

Is it safe to evaluate learner input with SymPy or math.js?

Only with care. SymPy's docs warn that sympify uses eval, and math.js recommends a separate worker or process with a timeout. Parse input into a restricted form and isolate the evaluator.

Has RAITHub built an AI math tutor?

Yes. RAITHub built PadhAI, an AI tutoring platform with a Socratic tutor, math verification, RAG and a 70/20/10 LLM router across PWA, WhatsApp and Telegram.

AI math tutormath verificationLLM math errorscomputer algebra systemSocratic tutorEdTechPadhAI

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.