Accuracy
How accurate is AI marking — and how would you know?
Everyone selling a marking tool has a percentage. Almost none of them say what it is a percentage of. Here is what the question actually means, why the honest answer is more complicated than a number, and how to find out for yourself in about twenty minutes.
Accurate against what?
There is no true mark sitting behind a piece of extended writing waiting to be discovered. There is a scheme, and there are examiners applying it, and they do not fully agree with one another. Studies of examiner agreement in essay-based subjects put full agreement on a given script at around half the time, with most of the disagreement landing a mark or two either side of a boundary.
So "97% accurate" is not a meaningful claim unless it says which marks it was compared against, how many scripts, in which subject, and how far apart the two marks were allowed to be before they counted as disagreeing. Measured generously enough, any tool is 97% accurate.
The comparison that matters to you is narrower and much more useful: does this tool land where I would have landed, on my scheme, on my students' work? That is answerable, and you are the only person who can answer it.
The twenty-minute test
Take five scripts you have already marked, ideally spread across the range rather than five middling ones. Do not tell the tool what you gave them. Paste your scheme, run them, and compare.
What you are looking for is not exact agreement. It is the shape of the disagreement:
- Scattered, within a mark or two. Normal. That is the same spread two human markers show. Usable.
- Consistently low across all five. The tool is deducting from an ideal answer instead of marking by best fit. This is the single most common failure and it will quietly deflate every class set you run.
- Consistently high. It is rewarding fluency. The strongest writer in the set will be overrated and the second-language student underrated.
- Fine on four, wildly off on one. Look at that one. Usually the answer is off-task, or it is the borderline you agonised over yourself.
A consistent offset is much less alarming than a random one, incidentally. A tool that comes in reliably two marks low is at least ranking the class correctly, and it is fixable by anchoring. A tool that is unpredictable is not.
Feedback quality is separate, and easier to judge
A mark can be right for the wrong reasons. Read the feedback on a script whose mark you agree with and check whether the reasoning is one you would put your name to.
The specific thing to look for is whether the feedback is anchored to the student's actual writing. A comment that quotes a phrase they wrote can be checked. A comment that says "develop your analysis further" cannot be, because it is true of every essay ever submitted and tells the student nothing about what to do on Tuesday.
What Paddle checks before a mark reaches you
Rather than a headline accuracy figure, here is what actually happens to every mark:
- Best fit is enforced in the instructions. The marking rules explicitly forbid constructing an ideal answer and deducting from it, and require the level to be set by the work as a whole.
- Band and score have to agree. If a criterion names a band and awards a score outside that band's range, the contradiction is caught and reconciled — every mark, not only uncertain ones. Where it has to be adjusted, the adjustment is flagged to you rather than made silently.
- Quoted evidence is verified against the script. Every criterion must quote the student's own words, and that quote is checked against the submitted text. If it cannot be found, the criterion is flagged. A tool that cites phrases the student never wrote is the failure mode nobody else tests for.
- Completeness is checked. If the criteria marked do not cover the paper's total, the mark fails rather than arriving looking plausible with a question missing.
- Uncertain marks get a second pass. Below a confidence threshold, a separate moderation step audits the mark against the scheme and the evidence cited.
- Confidence is reported to you. It answers one question: how likely is it that another experienced examiner would award the same mark? Low confidence is a prompt to look, not a defect.
None of that makes a mark right. It makes a wrong mark visible, which is the property you actually need from something whose output you are going to sign.
What we have not published
A validation study. Paddle has been tested against real marked scripts, including European Baccalaureate History papers marked by the teacher who set them, but that is a small sample and it would be dishonest to turn it into a percentage.
The study being run is straightforward: a set of scripts across two subjects, each marked by a teacher and by Paddle, recorded one row per criterion rather than one per script, keeping the sign of the difference so that a systematic bias in either direction shows up rather than cancelling out. When there is enough of it to mean something, it goes on this page, including the parts that do not flatter us.
Until then, the twenty-minute test above is a better guide than anyone's marketing number, including ours.
Run the test
Five scripts you have already marked. Your scheme. Twenty minutes. Then decide.
Try Paddle free →25 marks a month on the free plan. No card.