GCSE

Marking GCSE essays against your exam board scheme

A GCSE mark scheme is not a checklist. It is a set of levels with descriptors, and the skill is placing an answer in the right one rather than counting what is missing. That distinction is where most marking — human and automated — goes wrong.

Levels, not points

Extended-response GCSE questions are almost always marked by levels of response. The scheme describes what work at each level looks like, and each level covers a range of marks. Your job is two decisions in order: which level does this answer, as a whole, best fit; then where in that level's range does it sit.

That is "best fit", and it is deliberately not the same as awarding a point for each correct thing. An answer does not need everything a level describes to be placed in it. It needs the qualities that level describes, in the round, more than it has the qualities of the level below.

The opposite approach — imagining the perfect answer and deducting for each absence — produces marks that are consistently too low. It also produces feedback that is a list of everything the student is not, which no fifteen-year-old has ever acted on.

Assessment objectives

Most subjects split the marks across assessment objectives, and the weighting is not decoration. If a question is dominated by analysis and evaluation, a student who has written everything they know about the topic and evaluated none of it cannot reach the top levels no matter how much of it is correct.

The practical consequence when you are marking quickly: separate the objectives before you decide. Read once for knowledge, once for the analytical demand the question actually makes. Answers that feel strong on a first read are very often strong on one objective and thin on the one carrying most of the marks.

The mistakes that cost a class set half a grade

Marking to A-level standard

A Year 11 answer written in a lesson is not a sixth-form essay and should not be measured against one. At GCSE, naming an event correctly and explaining why it mattered is specific knowledge. Wanting more is not rigour; it is the wrong standard, applied uniformly, and it drags the whole set down.

Drifting across the pile

Script one and script thirty are not marked by the same person. You are more generous when you are fresh and harsher at eleven o'clock, or the other way round, and the ordering effect is real: a weak answer read after three strong ones scores lower than the same answer read first. This is the part of marking that is genuinely difficult to do well by hand, because it is not a knowledge problem.

Naming a level and awarding below it

Writing "Level 3" in the comment and putting a Level 2 number beside it happens constantly, usually out of caution. The student then reads Level 3 feedback next to a mark that contradicts it, and neither the number nor the comment can be defended. If you have named a level, the mark goes inside its range.

Credit for things the scheme does not credit

Neat handwriting, length, a confident tone, a tidy conclusion. Unless the scheme awards marks for accuracy of written communication, spelling and grammar are not worth marks either — and a student writing in their second language should be judged on their thinking, not on the surface of their prose.

Borderlines

Most real answers sit near a boundary, and rounding down every time is not the safe option. It is wrong in the same direction, systematically, across every student you teach.

Decide which level the answer genuinely best fits. If it sits squarely between two, a half mark is the correct answer to a genuine borderline — that is what it is for. What you should not do is mark down "to be safe" and then write feedback describing the level above.

Doing this for thirty scripts

Everything above is achievable on one essay. The difficulty is doing it identically thirty times in one evening, which is a consistency problem rather than a judgement problem. The same scheme applied the same way to all thirty, with no drift between the first and the last, is the specific thing a person cannot reliably do and a machine can.

That is the case for automating the first pass and keeping the second for yourself. You still read every mark; you are reading a proposal with the evidence quoted, rather than starting from a blank margin at nine in the evening.

Getting a usable mark out of any tool

How Paddle handles it

Paddle marks against the scheme you paste and nothing else — your criterion names, your maximum marks, your level descriptors. It marks by best fit rather than deducting from an ideal answer, marks to the level you state, and checks every score against the band named beside it, so a Level 3 comment cannot arrive with a Level 2 number. Source material you supply is excluded from any judgement of the student's writing.

Every criterion quotes the phrase from the student's own work the mark rests on, and every mark carries a confidence score so you know which ones to look at twice. You edit anything you disagree with and publish. Nothing reaches a student before that.

Test it on a script you have already marked

You know what that script is worth. Twenty minutes and five essays will tell you whether Paddle marks the way you do.

Try Paddle free →

25 marks a month on the free plan. No card.