Last year I was preparing my Year 4 class for their ISA writing exam, and I built a small machine. Every weekend, each student wrote a full essay for homework, in the exact format the exam demanded. On Monday they swapped papers and marked their partner's work against the same rubric the examiners would use — the same rubric every single week, until they knew it inside out. We quantified it. After each session, students read out their scores and I entered them into a spreadsheet that converted them into a percentage and lit up a traffic light: green for safe, red for danger. We ran this for weeks, and I watched the scores climb week by week. That cohort scored the highest results in that exam the school had ever seen.
So yes, I have a dog in this fight. But it is worth being honest about why peer-marking works, because it often doesn't.
The lazy-teacher critique
Some argue that peer-marking is teachers being lazy — outsourcing the professional feedback that students go to school to receive. The critique is valid, up to a point. It certainly saves time. But saving time is not the same as wasting the student's, and the research suggests something more interesting is going on.
Large-scale meta-analyses confirm that peer assessment, done properly, improves academic performance (Double, McGrane & Hopfenbeck, 2020; Li et al., 2020). Crucially, though, the benefit appears to be driven less by receiving a mark than by the cognitive work of producing a judgement. Evaluating a peer's essay means diagnosing gaps, applying success criteria and articulating solutions — higher-order skills that deepen the assessor's own understanding. Nicol, Thomson and Breslin (2014) found that students often gain more from giving feedback than from receiving it. The marker learns by marking. In my classroom that was the quiet engine of the whole system: after six weeks of marking the same rubric, my students could recite the exam criteria in their sleep, and their own essays showed it.
Where it goes wrong
The research is equally clear about the failure modes. Peer-assigned grades show only moderate reliability against expert judgement, and students consistently award higher marks to peers than teachers do (Power & Tanner, 2023). Without clear criteria, marking is buffeted by friendship, rivalry and uneven competence. Students know this, and many are anxious or sceptical about being graded by classmates (Fleckney, Thompson & Vaz-Serra, 2024). Used summatively — peer marks as formal grades — the practice degrades trust and produces unreliable data.
There is also a sobering statistic from Hattie and Timperley (2007): around 80 per cent of the verbal feedback students receive in class comes from other students, and much of it is incorrect. If you let peer-marking run unstructured, you are essentially amplifying the noisiest feedback channel in the room.
What the evidence says to do
The fixes are well established. Train the markers explicitly rather than assuming they can apply criteria by osmosis — trained raters are markedly more accurate and the learning gains larger (van Zundert, Sluijsmans & van Merriënboer, 2010; Li et al., 2020). Use a structured, analytical rubric: Fadillah and Ha (2024) found rubrics improved both scoring accuracy and learning outcomes, and reduced students' overconfidence in their own judgements. Keep it formative and low-stakes, so the mark is information rather than verdict. Anonymising the process can reduce social friction, though the evidence here is genuinely mixed (Panadero & Alqassab, 2019), so I would treat it as optional rather than essential.
Looking back, my Year 4 system ticked these boxes almost by accident. It was formative — nothing counted except the exam itself. The rubric was fixed, simple and quantified, which kept misinterpretation to a minimum. The weekly repetition was the rater training. Even the public scores, which can seem brutal, worked because the numbers were criterion-referenced rather than personal: everyone could see what the exam wanted, and the traffic light told you where you stood against it, not against each other's feelings. The trick is to channel that competitive energy respectfully, and to make sure the students in the red zone get clear instructions and enough support to climb out of it. A league table with no ladder is just cruelty.
So is peer-marking effective? Yes — on conditions. Formative, not summative. Trained markers, not hopeful ones. A rubric simple enough to be applied consistently and quantified honestly. Get those right and the marking session is not the teacher skiving; it is some of the most productive learning in the week. Get them wrong, and the critics are correct.