England · 2020
Exams were cancelled, so each grade was calculated from the school's past results and the teacher's ranking
The model produced a credible national distribution. Four days after it reached individual students, it was withdrawn.
Published 18 August 2026
What happened
On 18 March 2020 the Secretary of State announced that the summer exam series was cancelled, and that students would be awarded a grade based on an assessment of what they would have been most likely to achieve had the exams gone ahead. Ofqual asked every school and college for two things per student per subject: a centre assessment grade (CAG), the teacher's estimate of that grade, and a rank order of students within each grade. Over 5 million CAGs and rank positions were submitted through exam board portals in early June [1].
Ofqual's first analysis of those submissions found them optimistic. Awarded unchanged, they would have taken the share of A* grades at A level from 7.7% in 2019 to 13.9%, and the share of grades at B and above from 51.1% to 65% [1]. A ministerial direction of 3 April had required that overall results be, as far as possible, in line with previous years — but Ofqual's stated purpose for standardisation was not only that. It was also to stop a student's grade depending on how generously their own school had estimated: "as far as possible, ensure that a grade represents the same standard, irrespective of the school or college they attended" [1].
The method chosen was the direct centre-level performance (DCP) approach. For each centre in each subject it started from that centre's own distribution of grades over recent years, adjusted it for the prior attainment of the cohort that centre was entering this year (mean GCSE score, at A level), and then produced student grades by "overlaying the rank order provided by the centre onto each centre's predicted cumulative percentage grade distribution" [1]. The common description that the model ignored individual pupils is not quite right, and the report is precise about which part is true: "neither through the standardisation process applied this year nor the use of cohort level predictions in a typical year, does a student's individual prior attainment dictate their outcome in a subject. Measures of prior-attainment are only used to characterise and, therefore, predict for group relationships between students" [1]. One fact about the individual did decide their grade — their teacher's rank. Their own past results did not.
How much of a grade came from the model depended on how many students the centre entered in that subject. Ofqual set two thresholds on the harmonic mean of a centre's current and historical entry: below 5, the centre's CAGs became its distribution; between 5 and 15, the CAG and statistical distributions were blended on a taper; above 15, the statistical prediction stood alone. A centre with no historical data at all was treated as the limiting case and its students received their CAGs, to "avoid anomalous and potentially indefensible adjustments" [1]. The report's own subject tables show how unevenly that landed: in A level mathematics, 70.9% of 2,730 centres fell in the full statistical band and 9.1% in the small-cohort band, while in Latin 72.4% of 308 centres were small-cohort [1].
Calculated grades reached A level students on Thursday 13 August 2020, the day Ofqual published the interim report. Across 718,276 A level entries, 58.7% of CAGs were unchanged and 2.2% were raised; 39.1% were lowered — 35.6% by one grade, 3.3% by two and 0.2% by three or more [1]. Among centres with entries in all three years, the share of grade A and above rose at every centre type, but by 4.7 percentage points at independent schools, against 2.0 at secondary comprehensives, 1.7 at academies and 0.3 at sixth form, FE and tertiary colleges [1]. Ofqual's December breakdown by centre type showed 45.6% of A level entries adjusted down at FE establishments against 33.8% at independent schools, and attributed the spread to two causes together: how generous different centres' CAGs had been, and "the proportion of small cohorts within particular centres" [2]. Its later research put the entry-level picture the same way — the calculated grade matched the CAG for 59% of entries, was higher for just over 2%, and was lower for 39% [3].
Four days later the policy was withdrawn. On 17 August Ofqual's chair, Roger Taylor, opened a statement "We understand this has been a distressing time for students, who were awarded exam results last week for exams they never took", and announced "that students be awarded their centre assessment for this summer — that is, the grade their school or college estimated was the grade they would most likely have achieved in their exam — or the moderated grade, whichever is higher" [4]. Some 68% of A level candidates had at least one subject upgraded when final grades were issued, and 10.3% had been holding calculated grades that were, across their subjects, three grades or more below their CAGs [3]. All four UK regulators dropped their planned approach the same way, which the Office for Statistics Regulation later read as evidence of "inherent challenges in the task" rather than of one bad model [5].
Where the same method is ordinary, and defensible
Statistical moderation is not an emergency measure — it is how comparability between schools is normally maintained, and the OSR's review says so plainly: "the use of statistical models to support the setting and maintenance of standards at the cohort level is a common feature of awarding grades in a normal year". In a normal year, it adds, there is "extensive expert moderation and standardisation of grades, drawing on statistical analysis", a "well-established process, which may not be perfect, but which broadly speaking is accepted by the public as being an authoritative assessment of a student's performance" [5]. What changed in 2020 was the unit of the output: the models "were expected to predict a single grade for each individual on each course" [5]. The other half of the honest account is that the model was not the crude sorting device it was widely described as. Ofqual's own follow-up research, over 457,420 A level entries from 246,110 candidates, found that once centres and subject choices were taken into account there was "no evidence that candidates' socio-economic background, SEND status or the language spoken at home were associated with the likelihood of receiving a three-grade gap" [3]. The gap that the record documents runs between centres, not between demographic groups inside the same centre [2][3].
Where it burned
The model was fitted on one unit and used on another. A centre's historical distribution, adjusted for its cohort's prior attainment, is a reasonable statement about a group; it becomes a claim about one named 18-year-old only at the last step, when a rank order is laid over the curve and the boundaries fall where they fall [1]. The OSR found that "analyses were not performed at student level to test the impact on individuals", that the models "could not be developed and tested on the data that was used to run the model, nor could they be tested at the individual level", and that "there was limited human review of outputs of the models at an individual level prior to results day" [5]. The people who could have caught a wrong answer were available and not asked: "in the exams context, the teachers and lecturers know the students best" and were "best placed to identify potential issues with the results", but no country returned calculated grades to centres for checking before results day [5]. The teacher rankings the grades hung on carried uncertainty of their own that was never expressed as a range, because "a single grade estimate for each student had to be awarded" [5]. So an aggregate that held up arrived as one letter, with no interval around it and nobody who knew the student in the loop, on a single day, into university admissions [5].
The tell
When a score, grade or number about one person comes out of a model, ask what unit it was fitted and tested on — if the answer is a group that person belongs to, you are holding a statement about the group and an assumption about them.
Ofqual's model was built, by its own account, to characterise groups: prior attainment entered it only to "predict for group relationships between students", and what it produced was a distribution per school per subject. None of that is a mistake. The mistake was the last step, where a group estimate became an individual verdict with no interval around it and no one who knew the student reviewing it before it was sent. This is the ordinary shape of the failure rather than a pandemic-specific one — credit scores, risk models, fraud flags and performance rankings are all fitted on populations and then read as facts about a person. The question to ask is not whether the model is accurate, because it can be entirely accurate about the population and still wrong about you; it is whether anything about you, specifically, is allowed to change the answer. When the answer is nothing, the model has stopped being evidence and become the decision.
Share this case
The image has the link printed on it, so it still leads back here.
The check is a habit, and habits are trained. What to never hand off is about exactly this line: a model may produce the distribution, but a decision attached to one named person needs someone who knows that person to sign it. The 2020 grades are the clearest case of what happens when nobody does — the statistics were reviewed at length, the individual outputs were not.
Sources
Every source below was opened and read. Last verified 17 August 2026.
- [1] Awarding GCSE, AS, A level, advanced extension awards and extended project qualifications in summer 2020: interim report (Ofqual/20/6656/1) — Ofqual, 13 August 2020
- [2] Summer 2020 results analysis – GCSE, AS and A level: update to the interim report (Ofqual/20/6729) — Ofqual, 18 December 2020
- [3] Grading gaps in summer 2020: who was affected by differences between centre assessment grades and calculated grades? — Ofqual (GRADE joint initiative with Ofsted and the Department for Education), 29 July 2021
- [4] Statement from Roger Taylor, Chair, Ofqual — Ofqual, 17 August 2020
- [5] Ensuring statistical models command public confidence: learning lessons from the approach to developing models for awarding grades in the UK in 2020 — Office for Statistics Regulation, 2 March 2021