Noesa
← The Casebook

England · 2020

Exams were cancelled, so each grade was calculated from the school's past results and the teacher's ranking

The formula produced a believable national picture. Four days after it reached individual students, it was withdrawn.

Published 18 August 2026

What happened

On 18 March 2020 the Secretary of State announced that the summer exams were cancelled. Students would instead be given a grade based on what they would most likely have achieved had the exams gone ahead. Ofqual, England's exam regulator, asked every school and college for two things for each student in each subject: a centre assessment grade (CAG) — the teacher's estimate of the grade — and a ranking of students within each grade. Over 5 million grades and rank positions were submitted through the exam boards' websites in early June [1].

Ofqual's first look at those submissions found them generous. Awarded as they stood, they would have taken the share of A* grades at A level from 7.7% in 2019 to 13.9%, and the share of grades at B and above from 51.1% to 65% [1]. A ministerial direction of 3 April had required overall results to stay, as far as possible, in line with previous years. But Ofqual gave a second reason for adjusting them. It wanted to stop a student's grade depending on how generous their own school had been: "as far as possible, ensure that a grade represents the same standard, irrespective of the school or college they attended" [1].

The method chosen was called the direct centre-level performance (DCP) approach. For each school in each subject, it started from that school's own spread of grades over recent years. It adjusted that spread for how this year's students had done earlier (at A level, their average GCSE score). Then it produced each student's grade by "overlaying the rank order provided by the centre onto each centre's predicted cumulative percentage grade distribution" — in plain terms, by laying the teacher's ranking over the school's expected curve [1]. The common claim that the model ignored individual pupils is not quite right. The report is precise about which part is true. In its words: "neither through the standardisation process applied this year nor the use of cohort level predictions in a typical year, does a student's individual prior attainment dictate their outcome in a subject. Measures of prior-attainment are only used to characterise and, therefore, predict for group relationships between students" [1]. One fact about the individual did decide their grade — where their teacher ranked them. Their own past results did not.

How much of a grade came from the formula depended on how many students the school entered in that subject. Ofqual set two thresholds, using a particular kind of average (the harmonic mean) of the school's current and past entries. Below 5 students, the teachers' grades stood as they were. Between 5 and 15, the teachers' grades and the statistical prediction were blended on a sliding scale. Above 15, the statistical prediction stood alone. A school with no past data at all was treated as the extreme case and its students received their teachers' grades, to "avoid anomalous and potentially indefensible adjustments" [1]. The report's own subject tables show how unevenly that fell. In A level maths, 70.9% of 2,730 schools were in the full statistical band and 9.1% in the small-group band. In Latin, 72.4% of 308 schools were small-group [1].

Calculated grades reached A level students on Thursday 13 August 2020, the day Ofqual published its interim report. Across 718,276 A level entries, 58.7% of teachers' grades were unchanged and 2.2% were raised. 39.1% were lowered — 35.6% by one grade, 3.3% by two, and 0.2% by three or more [1]. Among schools with entries in all three years, the share of grade A and above rose at every type of school. But it rose by 4.7 percentage points at independent (fee-paying) schools, against 2.0 at comprehensives, 1.7 at academies, and 0.3 at sixth form, further education and tertiary colleges [1]. Ofqual's December breakdown showed 45.6% of A level entries adjusted down at further education colleges, against 33.8% at independent schools. It put the spread down to two causes together: how generous different schools' estimates had been, and "the proportion of small cohorts within particular centres" — small groups, whose grades were adjusted less or not at all [2]. Its later research described the picture the same way: the calculated grade matched the teacher's grade for 59% of entries, was higher for just over 2%, and lower for 39% [3].

Four days later the policy was withdrawn. On 17 August Ofqual's chair, Roger Taylor, opened a statement: "We understand this has been a distressing time for students, who were awarded exam results last week for exams they never took". He announced that students would receive their centre assessment grade — "the grade their school or college estimated was the grade they would most likely have achieved in their exam" — "or the moderated grade, whichever is higher" [4]. Some 68% of A level candidates had at least one subject upgraded when final grades were issued. 10.3% had been holding calculated grades that were, across their subjects, three grades or more below their teachers' estimates [3]. All four UK regulators dropped their planned approach in the same way. The Office for Statistics Regulation later read that as evidence of "inherent challenges in the task" rather than of one bad model [5].

Where the same method is ordinary, and defensible

Statistical moderation is not an emergency measure. It is how grades are normally kept comparable between schools. The Office for Statistics Regulation's review says so plainly: "the use of statistical models to support the setting and maintenance of standards at the cohort level is a common feature of awarding grades in a normal year". In a normal year, it adds, there is "extensive expert moderation and standardisation of grades, drawing on statistical analysis". It calls that a "well-established process, which may not be perfect, but which broadly speaking is accepted by the public as being an authoritative assessment of a student's performance" [5]. What changed in 2020 was what the formula was asked to produce. The models "were expected to predict a single grade for each individual on each course" [5]. The other half of the honest account is that the model was not the crude sorting machine it was widely described as. Ofqual's own follow-up research, over 457,420 A level entries from 246,110 candidates, took schools and subject choices into account. It found "no evidence that candidates' socio-economic background, SEND status or the language spoken at home were associated with the likelihood of receiving a three-grade gap" [3]. The gap the record documents runs between types of school, not between groups of students inside the same school [2][3].

Where it burned

The formula was built for one thing and used for another. A school's past results, adjusted for how its students did earlier, is a reasonable statement about a group. It becomes a claim about one named 18-year-old only at the last step, when a ranking is laid over the curve and the grade boundaries fall where they fall [1]. The statistics watchdog found that "analyses were not performed at student level to test the impact on individuals". It found the models "could not be developed and tested on the data that was used to run the model, nor could they be tested at the individual level". And it found "there was limited human review of outputs of the models at an individual level prior to results day" [5]. The people who could have caught a wrong answer were available and not asked. "In the exams context, the teachers and lecturers know the students best" and were "best placed to identify potential issues with the results" — but no country sent calculated grades back to schools for checking before results day [5]. The teacher rankings that the grades hung on carried uncertainty of their own, which was never shown as a range, because "a single grade estimate for each student had to be awarded" [5]. So a picture that held up for the country as a whole arrived as one letter per student, with no margin around it and nobody who knew the student in the loop, on a single day, straight into university admissions [5].

The tell

When a score, grade or number about one person comes out of a model, ask what it was built and tested on. If the answer is a group that person belongs to, you are holding a fact about the group and a guess about them.

Ofqual's formula was built, by its own account, to describe groups. Students' earlier results went in only to "predict for group relationships between students", and what came out was a spread of grades per school per subject. None of that is a mistake. The mistake was the last step, where a group estimate became an individual verdict — with no margin around it, and no one who knew the student checking it before it was sent. It is like guessing one child's height from the average of their class. This is the ordinary shape of the failure, not a pandemic one. Credit scores, risk models, fraud flags and performance rankings are all built on crowds and then read as facts about a person. So the question is not whether the model is accurate. It can be entirely accurate about the crowd and still wrong about you. The question is whether anything about you, specifically, is allowed to change the answer. When the answer is nothing, the model has stopped being evidence and become the decision.

Share this case

The image has the link printed on it, so it still leads back here.

The check is a habit, and habits are trained. What to never hand off is about exactly this line: a model may produce the spread, but a decision attached to one named person needs someone who knows that person to sign it. The 2020 grades are the clearest case of what happens when nobody does — the statistics were checked at length, the individual results were not.

Sources

Every source below was opened and read. Last verified 17 August 2026.

  1. [1] Awarding GCSE, AS, A level, advanced extension awards and extended project qualifications in summer 2020: interim report (Ofqual/20/6656/1) — Ofqual, 13 August 2020
  2. [2] Summer 2020 results analysis – GCSE, AS and A level: update to the interim report (Ofqual/20/6729) — Ofqual, 18 December 2020
  3. [3] Grading gaps in summer 2020: who was affected by differences between centre assessment grades and calculated grades? — Ofqual (GRADE joint initiative with Ofsted and the Department for Education), 29 July 2021
  4. [4] Statement from Roger Taylor, Chair, Ofqual — Ofqual, 17 August 2020
  5. [5] Ensuring statistical models command public confidence: learning lessons from the approach to developing models for awarding grades in the UK in 2020 — Office for Statistics Regulation, 2 March 2021