Parents book a seat. AI marks the paper. Results reach the inbox.
Thousands of bubbles, marked by hand, after every sitting.
A UK tuition provider prepares children for grammar-school entrance at 11+, and runs paper mock exams at multiple venues because the real exam is on paper. Between 52 and 125 candidates sit each mock, answering up to 180 questions across four sections. Every one of those sheets was marked by hand, and every marking slip landed in a parent’s report.
They needed the whole lifecycle in one system: parents booking and paying, personalised answer sheets printed, scans marked without a special scanner, scores standardised and ranked, and branded reports published to families. Then, in a second phase, the same question bank taken online.
Build started 8 April 2026. Live 9 June. First sitting 28 June.
One system for the whole exam lifecycle.
AI marking from ordinary scans
Each page is split into sections and read in parallel by a vision model. Reading focused crops rather than whole pages lifted accuracy from about 90% to about 98% in testing. No special forms, no dedicated hardware.
A guard against invention
The model must list the options it can actually see, and its answer has to be one of them. Anything else is rejected rather than scored. Reading runs at zero temperature, so the same scan always gives the same result.
Human review before scoring
Every batch lands in a review screen with uncertain sheets flagged. Admins correct any answer, or the detected name and ID, and finalise. Only finalised sheets are scored.
Question bank to printed paper
The question paper, the answer key and the OMR answer sheet are generated from the same data, so they cannot disagree — the class of error that quietly corrupts a whole cohort’s scores.
A parser that makes no AI calls
An earlier AI page-reader dropped the tail of multi-page comprehension passages. We replaced it with a deterministic parser over the PDF text layer: same output every time, verified across ten real papers with every section count correct.
Payments that don’t lose bookings
Stripe Checkout with three independent confirmation paths behind one idempotent handler, so a missed webhook heals itself instead of becoming the worst support ticket a centre can get.
An online exam engine
Section tabs, a question palette colour-coded by status, mark-for-review, countdown and auto-submit, section locking, and proctoring signals. Strict for mocks, relaxed for practice, on the same bank and scoring as the paper exams.
Analytics past the score
Question-by-question performance, percentile curves, time per question split into productive, unproductive and idle, and a live difficulty rating derived from how the cohort actually performed rather than what the author assumed.
Privacy built into the query
Leaderboards motivate children, so parents see their own child named and every other entry as “Student”. Anonymisation happens on the server: other children’s names never reach the browser.
From one question bank to a parent’s inbox.
No screenshots: the platform is live and holds children’s names, scores and parents’ payment records. This is the lifecycle it runs.
Two real incidents, both fixed without changing a single result.
Building the platform was half the job. It runs in Docker on one existing server the client already had — shared with their mail, DNS and other sites — listening only on the internal interface, with no new firewall ports opened and files kept on the same machine rather than a third-party bucket.
“The report says 8 is wrong when the answer is 8.”
Scoring was correct: the child had chosen −8. The report font had no glyph for the mathematical minus sign, so it printed as 8. A platform-wide scan found 37 of 800 questions with characters the font could not draw.
We embedded a full Unicode font and sourced Maths options from the bank rather than the scan. Before telling parents, we compared old and new output across 140,540 answer cells in all five exams with results at the time. No score changed.
“Results are published but parents got no email.”
The client suspected a limit around 100 children. There was none — an earlier sitting had emailed 106 families, and all 123 sheets had finalised, computed and published correctly.
The logs showed the real cause: the job queue’s memory store was full of records of finished jobs kept since launch, and being configured never to drop queued work, it refused new jobs — including the result emails.
Production access is read-only by default, and every change is explicitly approved before it runs. In both cases the cause was found from logs and data before anything was altered.
Nine weeks from first line of code to live.
Production figures read from the live database on 27 September 2026 — counts only, no personal data. Twelve sittings between late June and early September, at 52 to 125 candidates each.
Phase one — bookings, payments, paper marking, scoring and published results — runs in production today. Phase two — the online exam engine, practice tests, analytics, exam packs, bank-driven paper building and the teacher content portal — is built, tested end to end and handed over, awaiting the client’s go-live. We say which is which rather than blurring the two.