Founder & Lead Engineer, RAITHub
An adaptive learning engine estimates what each learner knows from their answers, then picks the next question to match. Most products need one of three methods: rules-based mastery (for example "4 of the last 5 correct"), Bayesian Knowledge Tracing, which tracks mastery per skill with 4 parameters, or item response theory, which scores ability against calibrated item difficulty. Start with rules and move up when your data allows.
If you would rather have it built for you, see how RAITHub would build this below.
This guide is for founders and product leads who keep hearing "adaptive" in pitch decks and need to know what it means in code, what data it needs, and which method fits the product they actually have. For the wider picture of building a learning product, read the EdTech pillar guide and RAITHub's EdTech software development page.
What is an adaptive learning engine, in plain terms?
It is two pieces of logic working in a loop. A learner model turns a history of answers into an estimate: "this learner probably knows fractions, probably does not know ratios yet". A selection policy uses that estimate to choose what comes next: another ratio question, a worked example, or a move to the next topic.
Everything people call "adaptive" is some version of that loop. The differences are in how the estimate is made (a counting rule, a probability per skill, or an ability score on a scale) and in what you have to feed it. A product with 200 hand-written questions and 50 pilot learners needs a different engine from a test provider with a calibrated bank of 10,000 items. Picking the heavy method too early is the most common way founders burn months here.
Which adaptive method should I use: rules, BKT or IRT?
Use the simplest method your data can support, and plan the data model so you can move up later without a rewrite.
| Method | What it estimates | Data it needs | Choose it when | Main weakness |
|---|---|---|---|---|
| Rules-based mastery | Mastered or not, per skill, by a counting rule | Attempts tagged with skills; no history needed to start | You are pre-launch or have a small pilot and teachers need to understand the rule | Ignores guessing and careless slips; thresholds are judgement calls |
| Bayesian Knowledge Tracing (BKT) | Probability that the learner knows each skill | Ordered attempt sequences per learner per skill, to fit 4 parameters per skill | Practice-heavy products where learning happens during use, such as homework and tutoring | Treats each skill separately; parameters need fitting and can be implausible |
| Item response theory (IRT) | A learner ability score, and difficulty (and discrimination) per item | Many responses per item to calibrate the bank | Placement tests and exams, where ability is roughly fixed during the test | Assumes ability does not change during the test; calibration needs volume |
| Elo-style rating | A running ability and difficulty estimate updated after each answer | Just the attempt stream; estimates improve as data arrives | You want IRT-like behaviour without a separate calibration study | Less principled than fitted IRT; needs tuning of the update size |
The Elo row is a practical middle path that researchers have documented in education; see Pelánek's review of Elo rating in adaptive educational systems.
How does rules-based mastery work, and is it enough?
Often, yes, for a first version. A rule such as "a skill is mastered after 4 correct answers in the last 5 attempts, with no hints" is simple to build, simple to test and, most importantly, simple for a teacher or parent to understand when they ask why the product moved a child on.
Rules get you three things early: a working loop, real attempt data, and an explainable baseline that any later model has to beat. Their weakness is that they treat a lucky guess and a real answer the same way. When teachers start reporting that learners are being moved on too soon, or held back on skills they clearly know, that is the signal that a probabilistic model is worth the cost.
How does Bayesian Knowledge Tracing work?
BKT treats each skill as a hidden switch, either "known" or "not known", and updates the probability that the switch is on after every answer. It was introduced by Corbett and Anderson in "Knowledge tracing: Modeling the acquisition of procedural knowledge" (User Modeling and User-Adapted Interaction, 1995).
The model has four parameters per skill, the same four that the open-source pyBKT library fits (it also supports a fifth, "forget", in extended variants):
- Prior: the chance a learner already knows the skill before practising.
- Learn (or transit): the chance they move from "not known" to "known" after a practice opportunity.
- Guess: the chance of a correct answer without knowing the skill.
- Slip: the chance of a wrong answer despite knowing it.
After each answer, Bayes' rule updates the estimate given whether the answer was right or wrong, then the learn parameter is applied for the practice opportunity just taken. When the probability crosses a threshold, the skill counts as mastered. A threshold of 0.95 is a widely used convention, but treat it as a product decision you test, not a constant.
What does a BKT update look like in TypeScript?
This is the whole update for one skill and one answer. It is deliberately small: the hard part of BKT is fitting the four parameters and tagging items correctly, not the arithmetic.
// Bayesian Knowledge Tracing: one update for one skill, one answer.
export interface BktParams {
pInit: number // P(L0): knows the skill before practice
pLearn: number // P(T): learns it on a practice opportunity
pGuess: number // P(G): correct without knowing
pSlip: number // P(S): wrong despite knowing
}
export function bktUpdate(pKnown: number, correct: boolean, p: BktParams): number {
// 1. Posterior: how likely they knew it, given this answer.
const posterior = correct
? (pKnown * (1 - p.pSlip)) /
(pKnown * (1 - p.pSlip) + (1 - pKnown) * p.pGuess)
: (pKnown * p.pSlip) /
(pKnown * p.pSlip + (1 - pKnown) * (1 - p.pGuess))
// 2. Transition: they may have learned it from this practice.
return posterior + (1 - posterior) * p.pLearn
}
export const MASTERY_THRESHOLD = 0.95
// Replay a learner's history for one skill, oldest attempt first.
export function masteryFromHistory(answers: boolean[], p: BktParams): number {
return answers.reduce((pKnown, correct) => bktUpdate(pKnown, correct, p), p.pInit)
}
// Example: plausible starting parameters for a new skill (tune with data).
const params: BktParams = { pInit: 0.2, pLearn: 0.15, pGuess: 0.2, pSlip: 0.1 }
const pKnown = masteryFromHistory([false, true, true, true], params)
const mastered = pKnown >= MASTERY_THRESHOLD
Three things to test before trusting it. First, the output always stays between 0 and 1. Second, a correct answer never lowers the estimate when guess plus slip is below 1; if a fitted skill has guess plus slip at or above 1, the model is degenerate and correct answers stop meaning "knows it", so reject those parameters. Third, replaying the stored attempt history reproduces the stored mastery value exactly, so a teacher's "why?" always has an answer.
How does item response theory (IRT) work at a practical level?
IRT puts learners and questions on the same scale. Each learner has an ability score; each item has a difficulty, and in richer models a discrimination. The model predicts the chance a learner of a given ability answers a given item correctly.
- 1PL (and the closely related Rasch model): items differ only in difficulty. As Columbia's Mailman School of Public Health explains, difficulty is the ability level at which a learner has a 50% chance of answering correctly, and discrimination is held fixed across items.
- 2PL: adds discrimination per item, which controls how sharply the chance of success rises with ability. A high-discrimination item separates learners just below and just above its difficulty well.
- 3PL: adds a guessing floor, useful for multiple choice.
In 2PL, the chance of a correct answer is 1 / (1 + e^(−a(θ − b))), where θ is ability, b is difficulty and a is discrimination. Computerised adaptive testing (CAT) uses this to pick, at each step, the item that tells you most about the learner's current ability estimate, which is how a placement test can finish in fewer questions than a fixed paper.
The practical catch is calibration. Item parameters are estimated from many learners answering each item, so IRT suits a mature item bank or a test you can pilot at scale. For a new product with a small bank, that volume usually is not there yet, and the estimates will be noisy. Also note what IRT assumes: ability is stable during the test. That fits placement and exams. It fits poorly in a tutoring session where the whole point is that the learner improves as they go, which is the situation BKT was designed for.
What data do I need to build adaptive assessment?
An attempt log and a well-tagged item bank. Every method above reads from the same raw material, so get this right first and the method becomes a swappable component.
| Data | What to store | Why it matters |
|---|---|---|
| Attempt events | Learner, item, timestamp, correct or not, attempt number, hints used, response time | Append-only history is what every model replays; never overwrite it |
| Item-to-skill tags | Which skills each item tests (often called a Q-matrix) | A mis-tagged item silently corrupts every estimate for that skill |
| Skill graph | Prerequisites between skills | Lets the selection policy avoid jumping ahead |
| Model version | Which parameters and rules produced each mastery value | So you can explain a past decision after the model changes |
| Item metadata | Format, difficulty label, retired or live | Seeds cold start and lets you pull bad items without losing their history |
Because the people answering are often children, collect only what the model needs and treat the attempt log as personal data. The EdTech pillar covers the privacy rules that commonly apply. That is general information; confirm with your adviser which rules apply to your product.
How do I handle cold start for a new adaptive product?
Cold start is the period when you have no data on a new learner, a new item, or the whole product. Handle each separately:
- New learner. Start from a default prior, or run a short placement quiz. Let a teacher or the learner set a starting level when that is honest about what they know.
- New item. Seed its difficulty from an author's label (easy, medium, hard), mark it as uncalibrated, and serve it alongside calibrated items until it has enough responses.
- New product. Ship rules-based mastery, log everything, and fit BKT or IRT parameters once you have real sequences. Shared default parameters across similar skills are a sensible bridge until each skill has its own data.
The point is that the first release does not need a fitted model. It needs a data model that will let you fit one later.
Should I buy an adaptive engine, use a template, or build my own?
| Option | Examples | Cost signal | Choose this when |
|---|---|---|---|
| Off-the-shelf assessment API | Learnosity (authoring, delivery, scoring and reporting via APIs) | Quote-based; Learnosity says pricing scales with usage and monthly active users | Assessment is a feature of your product, not the product, and you can live with the vendor's model and pricing |
| LMS or no-code content tool | Moodle quizzes (open source); H5P.com | Moodle is free to self-host; H5P.com's own calculator shows one Premium configuration at 1,560 USD a year | You need quizzes and retries fast. Note Moodle's "adaptive mode" means multiple tries with penalties, not per-learner question selection |
| Open-source library plus your code | pyBKT (MIT licence) to fit BKT parameters offline | Free library; your engineering time | You have a developer comfortable with Python and data, and want to own the model |
| Custom build | Your own attempt log, learner model and selection policy | Fixed-scope project | Adaptivity is your product's core value and you need it explainable, testable and portable |
How long does it take to build it yourself, and what is the main risk?
As a rough guide from engineering experience rather than a published benchmark: rules-based mastery on a clean attempt log is 1–2 weeks for a developer who knows your stack; BKT with offline parameter fitting and a replayable history is 3–6 weeks; calibrated IRT with a CAT selection policy is longer and needs the response volume first.
The main risk is not the maths. It is mis-tagged items, overwritten progress values and untested update logic, which together produce mastery numbers nobody can explain. A scoring bug in a learning product is a fairness problem: a learner held back or pushed on for no reason. That is why the update function deserves the strictest tests in the codebase, the same way any LLM component does (see how to test LLM features).
Where does an AI tutor fit with knowledge tracing?
They solve different problems. The learner model decides what the learner should work on next; an AI tutor helps while they work on it. A sound design keeps mastery decisions in deterministic, tested code and lets the language model explain, hint and question, never silently decide that a skill is mastered. RAITHub's guide on how to build an AI tutor covers the tutor side.
Why RAITHub for this
- EdTech we have shipped. RAITHub built PadhAI, an AI tutoring platform: 11 services (9 Node/TypeScript, 2 Python), a 70/20/10 LLM router to keep model cost per learner viable, a Socratic tutor with math verification and RAG, and the same experience on a PWA, WhatsApp and Telegram, with 9 payment gateways. This post does not claim PadhAI uses BKT or IRT specifically; the methods here are engineering guidance.
- QA-first. Learner-model logic is exactly the code that needs property tests and replay tests. RAITHub's own marketplace, TheSkinProof (the founder's own venture), runs 750+ automated tests; PropDesk has 1,024.
- Polyglot where it helps. Product logic in TypeScript, parameter fitting in Python, as PadhAI's split already does.
When you don't need us
- You need quizzes with retries inside a course: Moodle or H5P will do it.
- Assessment is a side feature and a vendor's API and pricing work for you.
- You need psychometric validation of a high-stakes exam. That is a job for a psychometrician; RAITHub can build the platform around their model, but does not validate tests.
- You want the learner model delivered in a native mobile app. RAITHub builds web apps and PWAs only.
How RAITHub would build this
- Scope: an append-only attempt log and item bank with skill tags; rules-based mastery behind a learner-model interface; a BKT module with parameters fitted offline and versioned; a teacher view that explains each mastery decision from real attempts; and admin tools to retire or re-tag items.
- Timeline: 4–6 weeks at fixed scope for an MVP like this; 6–12 weeks if it is a backend or API that other products call.
- You receive: automated tests and CI (including replay and property tests for the learner model), handover docs and runbooks, and full IP under NDA.
- Next step: the free 15-minute audit, then a written fixed quote. Book the EdTech audit, or see the MVP development service and try the MVP cost estimator first.
Frequently asked questions
What is the difference between IRT and knowledge tracing?
IRT estimates a stable ability score and per-item difficulty, which suits placement tests and exams. Knowledge tracing, such as BKT, estimates the probability a learner knows each skill and expects that to change as they practise, which suits tutoring and homework.
How many parameters does Bayesian Knowledge Tracing use?
Four per skill in the standard model: prior knowledge, learn rate, guess and slip. Extended variants, such as those in the pyBKT library, add a forget parameter.
Do I need machine learning to make my product adaptive?
No. A rules-based mastery threshold on a well-tagged item bank is a legitimate first version, and it gives you the data to fit BKT or IRT later.
How much data does IRT need?
Enough responses per item to estimate its parameters reliably, which a new product with a small pilot usually does not have. Start with rules or BKT and calibrate items once the response volume exists.
What mastery threshold should I use with BKT?
0.95 is a widely used convention, but it is a product decision. Test it against teacher judgement and later performance on your own learners.
Can an AI tutor replace knowledge tracing?
It should not. Keep mastery decisions in deterministic, tested code that can be explained, and let the language model explain and hint.
How long does it take to build an adaptive learning MVP?
RAITHub scopes a focused adaptive learning MVP at 4–6 weeks fixed scope, after a free 15-minute audit and a written quote.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.