For teachers

AI essay marking for teachers

A class set takes an evening. By the time it comes back, the lesson it belonged to is three weeks gone and the student reads the number, not the comment. AI marking is worth having only if it fixes that particular problem — and only if you still decide every mark.

Most of what is written about AI and marking argues about whether a machine can judge an essay. That is the wrong argument to be having in a school. The interesting question is narrower: can a tool apply your mark scheme, to all thirty scripts, to the same standard, fast enough that the work comes back while it still matters — and leave the actual decision with you?

This page is about what that looks like in practice, what it cannot do, and what to check before you let any tool near a set of real student work.

Marking against your scheme, not a generic rubric

The first thing to check about any marking tool is whose standard it is applying.

A lot of them ship a library of built-in rubrics: pick "GCSE English Literature", get a mark. That is fast and it is useless, because the scheme your department actually uses is a specific document with specific band descriptors, and it is the one your students have been taught against. A mark produced against something else is not a mark your class can act on, and it is not one you can defend to a parent.

The alternative is to paste your own scheme in — the real one, the wording your department agreed — and have the tool mark against that and nothing else. Criterion names come from your document. Maximum marks come from your document. Bands come from your document. If your scheme says "Level 3 (11–15)", that is the label the student sees, and the score has to sit inside it.

That constraint matters more than it sounds. Marking is mostly the discipline of not inventing criteria, and a model left to its own devices will happily reward things your scheme does not credit — presentation, length, a confident tone.

Where automated marking usually goes wrong

Three failure modes account for most of it, and they are worth knowing about whether or not you ever use a tool, because they are also how you spot a bad one in about four minutes.

Marking down from an ideal answer

The most common one. The tool constructs the perfect response in its head, then deducts for everything the student did not do. Every mark comes out low, uniformly, and a teacher who trusts it hands back a class set that would fail moderation.

Real examiners mark by best fit: find the band the answer as a whole sits in, then place it within that band. An answer does not need every available point to be a Level 4. It needs the qualities Level 4 describes.

Marking to the wrong level

A Year 11 practice essay written in forty minutes is not an undergraduate seminar paper. A tool that demands treaty article numbers and historiographical debate from a sixteen-year-old is not being rigorous, it is marking the wrong exam. Naming the level — GCSE, A-level, S6, Year 10 — should change the standard applied, not just the wording of the feedback.

Rewarding writing the student did not do

Source-based questions hand the student a passage. If a tool reads a well-written extract inside the answer and credits the student for the prose, the mark is wrong and the strongest praise goes to whoever quoted most. Selection and use of a quote is creditable. The quality of the quote is not.

The part that has to stay with the teacher

A mark that goes straight to a student is a grade issued by a machine. A mark that arrives in front of a teacher, who reads it, changes what they disagree with, and presses publish, is a first draft done by a machine. Those are different products with the same technology underneath, and only one of them belongs in a classroom.

Practically, that means three things are non-negotiable:

The published mark being yours is also what makes the disagreement useful: where you consistently move a mark, that is a signal about the scheme or about the tool, and it is worth watching.

What it does not replace

Being clear about this is not modesty, it is the difference between a tool that survives contact with a department and one that gets banned in a term.

Before you use any tool on real work

That last one is the only evaluation that counts. You already know what those five scripts are worth. If the tool agrees with you on four of them and the fifth is a genuine borderline, it is doing the job. If it comes in three marks low on every single one, it is deducting from an ideal answer and you have just saved yourself a term of quietly deflated grades.

How Paddle does it

Paddle is built on exactly the constraints above. You paste your own mark scheme. You share one link with the class; students submit without creating an account and are asked for a first name only. Paddle proposes a mark and criterion-by-criterion feedback, each criterion quoting the phrase from the student's own work it rests on, and gives you a confidence score so you know which marks are worth a second look. You review, edit, and publish. Nothing reaches a student before that.

Student work is stored in the EU and is never used to train AI models. It is for formative work, not for final grades — which is also what the Department for Education's position on AI marking assumes.

Mark one real script

Paste a scheme, paste a script you have already marked, and see whether Paddle agrees with you. That is the only demo worth having.

Try Paddle free →

25 marks a month on the free plan. No card.