Founder & Lead Engineer, RAITHub
A quiz and assessment engine has five parts: an item bank, pools that randomise what each learner sees, an attempt runner, a scoring function, and item analytics. Scoring and analytics are where builds go wrong. Partial credit must stop learners ticking every box, and every question needs a discrimination index; the University of Washington rates any item below 0.10 as poor.
If you would rather have the engine built for you, see how RAITHub would build this below.
This guide is for EdTech founders, training platforms and product teams adding assessment to a learning product. It covers the design choices in the order you will meet them, with a PostgreSQL schema and a TypeScript scoring function you can adapt. Sources are the University of Washington's item analysis guide, Moodle's quiz documentation, 1EdTech and vendor pages, checked on 2 October 2026. If you are still deciding whether to build at all, read custom LMS vs Moodle first.
What item types should a quiz engine support first?
Start with the types you can score automatically and reliably, and add the rest when a customer needs them. An item is one question, with its answer key, points and feedback.
| Item type | Auto-scored? | Partial credit? | Build effort | Main trap |
|---|---|---|---|---|
| Single choice | Yes | No | Low | Option order bias if you never shuffle |
| Multiple response (tick all that apply) | Yes | Yes | Low | Ticking every option earns full marks unless wrong ticks cost something |
| True or false | Yes | No | Low | A 50% guess rate; negative marking matters most here |
| Numeric | Yes | Optional, by tolerance band | Medium | Units, rounding and locale decimal separators |
| Short text | Partly | Rarely | Medium | Spelling and synonyms; plan for manual review |
| Ordering and matching | Yes | Yes, per correct pair or position | Medium | Defining what "half right" means before launch |
| Maths expressions | Yes, with an expression checker | Possible | High | Equivalent answers (2x+2 and 2(x+1)) marked wrong by string matching |
| Essay and file upload | No | Rubric-based | Medium | A marking workflow, not just a field |
For a sense of a baseline, Google Forms quizzes auto-grade "Short answer, Multiple choice, Checkboxes, Dropdown, Multiple choice grid, Checkbox grid" (Google Docs Editors Help). If your needs stop there, you may not need an engine at all. If you must move items between systems, the 1EdTech QTI standard is "a way of packaging and moving tests and questions from one application to another" (1EdTech QTI); keep your own item format close enough to map onto it later.
How do question pools and randomisation work?
An assessment is a list of sections, and each section draws N items from a pool. Randomisation then works at two levels: which items are drawn, and the order of items and options.
- Store a seed per attempt. Generate the draw and the shuffle from a seed saved on the attempt, and freeze the result in an attempt_item table when the attempt starts. A page reload, a second device or an appeal must all show the same paper.
- Balance the pools. If a pool mixes easy and hard items, two learners can sit very different tests. Tag items by topic and difficulty and draw per tag, so every paper has the same shape.
- Freeze items once used. Never edit an item that has responses. An edit creates a new item row that supersedes the old one, so past scores still match the question the learner actually saw.
- Do not shuffle everything. Options such as "All of the above" or a numeric scale must keep their order; make shuffling a per-item flag.
How should scoring handle partial credit and negative marking?
Pick a rule per assessment, write it down for learners, and make the code match it exactly. Three rules cover most needs:
| Rule | How it scores a multiple-response item | Use it when |
|---|---|---|
| All or nothing | Full points only for exactly the correct set | Certification-style tests where half-right is wrong |
| Partial credit | Credit for each correct tick, minus credit for each wrong tick, never below zero | Learning and practice, where progress matters |
| Partial with negative marking | As partial, but a wrong answer can score below zero, down to a stated floor | High-stakes multiple choice where guessing must not pay |
The tick-everything problem is real. Moodle's documentation warns that if you allow multiple answers with more than one correct choice "and do not use a negative grade percentage for wrong answers, the students can simply tick all choices and get the full grade" (Moodle multiple choice question type). Two more rules prevent most disputes: an unanswered item always scores zero, never a penalty, and scores are rounded once, at item level, the same way everywhere.
A minimal scoring function in TypeScript, covering choice and numeric items under all three rules:
type Rule = 'all_or_nothing' | 'partial' | 'partial_negative'
interface ChoiceItem { kind: 'choice'; points: number; correct: Set<string>; optionCount: number }
interface NumericItem { kind: 'numeric'; points: number; value: number; tolerance: number }
type Item = ChoiceItem | NumericItem
type Answer =
| { kind: 'choice'; selected: string[] }
| { kind: 'numeric'; value: number }
| null
/** negativeFloor: the most a wrong answer can cost, as a fraction of points (e.g. 0.25). */
export function scoreItem(item: Item, answer: Answer, rule: Rule, negativeFloor = 0): number {
if (answer === null) return 0 // unanswered is never penalised
const floor = rule === 'partial_negative' ? -negativeFloor : 0
let fraction: number
if (item.kind === 'numeric') {
if (answer.kind !== 'numeric' || !Number.isFinite(answer.value)) return 0
fraction = Math.abs(answer.value - item.value) <= item.tolerance ? 1 : floor
} else {
if (answer.kind !== 'choice') return 0
const selected = new Set(answer.selected) // ignore duplicate ids
let hits = 0
for (const id of selected) if (item.correct.has(id)) hits++
const wrong = selected.size - hits
const wrongOptions = item.optionCount - item.correct.size
if (rule === 'all_or_nothing') {
fraction = hits === item.correct.size && wrong === 0 ? 1 : 0
} else {
const raw = hits / item.correct.size - (wrongOptions > 0 ? wrong / wrongOptions : 0)
fraction = Math.min(1, Math.max(floor, raw))
}
}
return Math.round(fraction * item.points * 100) / 100
}
Check it against the edge cases before anything else: a learner who ticks all four options of an item with two correct answers scores 1 − 2/2 = 0; a wrong single choice from four options under negative marking with a floor of 0.25 scores −0.25 of the item's points. Those cases belong in your test suite, along with empty answers and duplicate ids.
How should attempts, timers and feedback work?
- Attempts. Store max_attempts on the assessment and number each attempt; a unique constraint on assessment, learner and attempt number stops a double-click creating two. Decide which attempt counts: highest, latest, first or average.
- Timers. The server owns the clock. Save started_at, compute the deadline on the server, and reject or auto-submit responses after it. A client-side countdown is only a display.
- Autosave. Save each response as it is given, so a dropped connection loses nothing. One row per attempt and item, upserted.
- Feedback timing. Immediate feedback suits practice; feedback after the attempt closes suits graded tests, because early feedback leaks the answer key to the next learner. Google Forms makes the same split, releasing grades "Immediately after each submission" or "Later, after manual review" (Google Docs Editors Help).
- Regrading. When a key is found wrong, rescore stored responses with the corrected key and keep an audit row of the old and new score.
What does the database schema look like?
A PostgreSQL starting point. Items are immutable once used, attempts freeze their paper, and responses keep the raw answer as JSON so they can be rescored.
CREATE TABLE item (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
type text NOT NULL CHECK (type IN ('single_choice','multi_choice','numeric','short_text','ordering')),
stem text NOT NULL,
points numeric(6,2) NOT NULL DEFAULT 1 CHECK (points > 0),
answer_key jsonb NOT NULL, -- correct option ids, or value and tolerance
feedback text,
supersedes uuid REFERENCES item(id), -- edits create a new row
created_at timestamptz NOT NULL DEFAULT now()
);
CREATE TABLE item_option (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
item_id uuid NOT NULL REFERENCES item(id),
body text NOT NULL,
position int NOT NULL,
shuffle boolean NOT NULL DEFAULT true
);
CREATE TABLE pool (id uuid PRIMARY KEY DEFAULT gen_random_uuid(), name text NOT NULL);
CREATE TABLE pool_item (pool_id uuid REFERENCES pool(id), item_id uuid REFERENCES item(id),
PRIMARY KEY (pool_id, item_id));
CREATE TABLE assessment (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
title text NOT NULL,
max_attempts int NOT NULL DEFAULT 1 CHECK (max_attempts > 0),
time_limit_s int CHECK (time_limit_s > 0),
scoring_rule text NOT NULL CHECK (scoring_rule IN ('all_or_nothing','partial','partial_negative')),
negative_floor numeric(4,2) NOT NULL DEFAULT 0 CHECK (negative_floor BETWEEN 0 AND 1),
feedback_mode text NOT NULL CHECK (feedback_mode IN ('immediate','after_submit','after_close'))
);
CREATE TABLE assessment_section (
assessment_id uuid REFERENCES assessment(id),
position int NOT NULL,
pool_id uuid NOT NULL REFERENCES pool(id),
draw_count int NOT NULL CHECK (draw_count > 0),
PRIMARY KEY (assessment_id, position)
);
CREATE TABLE attempt (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
assessment_id uuid NOT NULL REFERENCES assessment(id),
learner_id uuid NOT NULL,
attempt_no int NOT NULL,
seed bigint NOT NULL,
started_at timestamptz NOT NULL DEFAULT now(),
submitted_at timestamptz,
score numeric(8,2),
UNIQUE (assessment_id, learner_id, attempt_no)
);
CREATE TABLE attempt_item (
attempt_id uuid REFERENCES attempt(id),
item_id uuid REFERENCES item(id),
position int NOT NULL,
PRIMARY KEY (attempt_id, item_id)
);
CREATE TABLE response (
attempt_id uuid,
item_id uuid,
answer jsonb NOT NULL,
score numeric(6,2),
answered_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (attempt_id, item_id),
FOREIGN KEY (attempt_id, item_id) REFERENCES attempt_item (attempt_id, item_id)
);
The composite foreign key on response is the quiet hero: a learner cannot answer an item that was never on their paper. If many organisations share the engine, add a tenant column and row-level security, as in PostgreSQL row-level security for multi-tenant apps.
How do you measure item difficulty and the discrimination index?
Two numbers per item tell you whether a question is doing its job.
- Difficulty (also called the facility index or p-value). For a one-point item it "is simply the percentage of students who answer an item correctly"; the University of Washington treats 85% or above as easy, 51–84% as moderate and 50% or below as hard (UW item analysis).
- Discrimination: the correlation between the item score and the learner's score on the rest of the test. The same guide rates above 0.30 as good, 0.10–0.30 as fair and below 0.10 as poor. A negative value usually means a wrong answer key or a confusing question. Moodle computes the same measure and notes that "unless the facility index is 50%, it is impossible for the discrimination index to be 100%" (Moodle quiz statistics).
Both can come straight out of PostgreSQL, which has a built-in corr() aggregate. This uses first attempts only, and skips items with too few responses to mean anything:
WITH scored AS (
SELECT r.item_id,
r.score / i.points AS item_frac,
a.score - r.score AS rest_score
FROM response r
JOIN item i ON i.id = r.item_id
JOIN attempt a ON a.id = r.attempt_id
WHERE a.submitted_at IS NOT NULL AND a.attempt_no = 1
)
SELECT item_id,
count(*) AS responses,
round(avg(item_frac), 2) AS difficulty,
round(corr(item_frac, rest_score)::numeric, 2) AS discrimination
FROM scored
GROUP BY item_id
HAVING count(*) >= 30
ORDER BY discrimination NULLS FIRST;
With pools, each learner's "rest of test" is a different set of items, so treat these figures as a screening tool: they flag items to review, not items to delete automatically. Add distractor analysis (how often each wrong option is chosen) next; an option nobody picks is not doing any work.
Buy, build or hire?
| Option | Example | Choose this when |
|---|---|---|
| Form tool with quizzes | Google Forms: point values, answer feedback and auto-grading for choice, checkbox and short-answer items | Low-stakes quizzes for a class or team, with no pools or analytics needed |
| Hosted testing service | ClassMarker: subscription plans or credit packs, a 30-day free trial; current prices on its pricing page | You need question banks, random questions and certificates, and the tests are not your product |
| Open-source LMS quiz | Moodle: no licence fee; positive and negative grade percentages per option, shuffling, and item statistics | Assessment lives inside standard courses and someone can run Moodle |
| Custom engine | Your schema, your scoring rules, your analytics, in your product | Assessment is part of what customers pay for: adaptive practice, exam prep, AI tutoring or a scoring model no tool supports |
Doing it yourself: an experienced full-stack developer can build choice and numeric items, pools, attempts, scoring and a basic report in roughly 3–6 weeks; that is an engineering estimate, not a vendor figure. The main risk is scoring errors found after results are released. Without stored raw answers and a tested regrade path, a wrong key means manual fixes and unhappy learners.
Why RAITHub for this
RAITHub built PadhAI, an AI tutoring platform: 11 services (9 Node/TypeScript, 2 Python), a Socratic tutor with math verification and retrieval-augmented generation, a 70/20/10 LLM router, delivery across a PWA, WhatsApp and Telegram, and 9 payment gateways. Math verification is the hard edge of assessment, where a correct answer can be written many ways, and the Socratic tutor is practice and feedback built into the learning loop. The design behind it is in how to build an AI tutor. RAITHub has not published a standalone quiz engine with item analytics; that part would be scoped as new work. What carries over is the QA-first approach: scoring rules are written as tests before they are written as code, through QA and test automation as standard.
When you don't need us
- Your quizzes are low-stakes and few. A form tool is enough.
- You need banks, random draws and certificates for staff or candidates. A hosted testing service will be live this week.
- Your assessment sits inside standard courses. Moodle's quiz already does pools, negative grades and statistics.
- You need psychometric validation or a certified exam programme. That needs a psychometrician and an accredited body, not only software.
How RAITHub would build this
- Scope: an item bank with versioned items, pools and tags, and the item types your first course needs.
- An attempt runner with seeded randomisation, a server-owned timer, autosave and attempt limits.
- A scoring module covering all-or-nothing, partial credit and negative marking, with a regrade path and an audit trail.
- Item analytics: difficulty, discrimination and distractor reports, with CSV export.
- Optional: maths expression checking or AI-assisted feedback, routed and tested as in PadhAI.
Timeline: as a fixed-scope MVP, 4–6 weeks; added as a backend and API to an existing product, 6–12 weeks depending on item types and integrations.
You receive: automated tests and CI, including a scoring test suite of edge cases, handover docs and runbooks, and full IP under an NDA. RAITHub signs NDAs and DPAs and works inside your controls; production learner data stays in your own cloud account, and development uses synthetic data. Builds run under SaaS development as fixed scope or a dedicated team.
Next step: book the free 15-minute technical audit with a sample of your questions and your scoring rules, and you will get a written fixed quote. The EdTech page and the EdTech software development guide cover the wider product.
University of Washington, Moodle, 1EdTech, Google and ClassMarker sources checked on 2 October 2026.
Frequently asked questions
What is a discrimination index in a quiz?
It measures whether learners who do well on the rest of the test also get this question right, as a correlation. The University of Washington rates above 0.30 as good and below 0.10 as poor; a negative value often means a wrong key.
How do you stop learners ticking every option in a multiple-answer question?
Make wrong ticks cost credit. With partial scoring that subtracts for each incorrect option, ticking everything scores zero. Moodle's documentation warns that without negative grades, ticking all choices earns the full grade.
Should a quiz engine use negative marking?
Only for high-stakes multiple choice where guessing must not pay, and only with the rule and the floor stated to learners in advance. For practice and learning, partial credit without penalties is usually fairer.
How do you randomise questions fairly?
Draw items from pools tagged by topic and difficulty, so every paper has the same shape. Save a seed per attempt and freeze the drawn items when the attempt starts, so reloads and appeals show the same paper.
How long does it take to build a quiz engine?
Choice and numeric items, pools, attempts, scoring and basic reports take an experienced developer about 3–6 weeks. RAITHub's fixed-scope MVP range is 4–6 weeks; analytics, maths checking and integrations add time.
Can I import questions from another system?
Often, if both sides support 1EdTech QTI, the standard for moving tests and questions between applications. Otherwise plan a one-off import script and check every answer key after import.
Has RAITHub built an assessment engine?
RAITHub built PadhAI, an AI tutoring platform with a Socratic tutor and math verification. A standalone quiz engine with item analytics would be scoped as new work.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.