Accuracy

How accurate is AI marking — and how would you know?

Everyone selling a marking tool has a percentage. Almost none of them say what it is a percentage of. Here is what the question actually means, why the honest answer is more complicated than a number, and how to find out for yourself in about twenty minutes.

Accurate against what?

There is no true mark sitting behind a piece of extended writing waiting to be discovered. There is a scheme, and there are examiners applying it, and they do not fully agree with one another. Studies of examiner agreement in essay-based subjects put full agreement on a given script at around half the time, with most of the disagreement landing a mark or two either side of a boundary.

So "97% accurate" is not a meaningful claim unless it says which marks it was compared against, how many scripts, in which subject, and how far apart the two marks were allowed to be before they counted as disagreeing. Measured generously enough, any tool is 97% accurate.

The comparison that matters to you is narrower and much more useful: does this tool land where I would have landed, on my scheme, on my students' work? That is answerable, and you are the only person who can answer it.

The twenty-minute test

Take five scripts you have already marked, ideally spread across the range rather than five middling ones. Do not tell the tool what you gave them. Paste your scheme, run them, and compare.

What you are looking for is not exact agreement. It is the shape of the disagreement:

A consistent offset is much less alarming than a random one, incidentally. A tool that comes in reliably two marks low is at least ranking the class correctly, and it is fixable by anchoring. A tool that is unpredictable is not.

Feedback quality is separate, and easier to judge

A mark can be right for the wrong reasons. Read the feedback on a script whose mark you agree with and check whether the reasoning is one you would put your name to.

The specific thing to look for is whether the feedback is anchored to the student's actual writing. A comment that quotes a phrase they wrote can be checked. A comment that says "develop your analysis further" cannot be, because it is true of every essay ever submitted and tells the student nothing about what to do on Tuesday.

What Paddle checks before a mark reaches you

Rather than a headline accuracy figure, here is what actually happens to every mark:

None of that makes a mark right. It makes a wrong mark visible, which is the property you actually need from something whose output you are going to sign.

What we have not published

A validation study. Paddle has been tested against real marked scripts, including European Baccalaureate History papers marked by the teacher who set them, but that is a small sample and it would be dishonest to turn it into a percentage.

The study being run is straightforward: a set of scripts across two subjects, each marked by a teacher and by Paddle, recorded one row per criterion rather than one per script, keeping the sign of the difference so that a systematic bias in either direction shows up rather than cancelling out. When there is enough of it to mean something, it goes on this page, including the parts that do not flatter us.

Until then, the twenty-minute test above is a better guide than anyone's marketing number, including ours.

Run the test

Five scripts you have already marked. Your scheme. Twenty minutes. Then decide.

Try Paddle free →

25 marks a month on the free plan. No card.