What happened
In the UK, women are offered a routine breast X-ray — a mammogram — to catch cancer early, and every scan is checked by two specialist doctors, one after the other. In 2024 the health board for north-east Scotland, NHS Grampian, working with the University of Aberdeen, reported a test of a computer program called Mia, made by a company called Kheiron Medical Technologies. The program was run over the scans of more than ten thousand women — scans that both doctors had already checked and cleared [1].
The program pointed at a small number of scans — reported as eleven — and said, in effect, look at this one again. They turned out to show cancers that both doctors had missed. The flagged scans went back to the doctors, who made every decision about whether to call a woman back for more tests; the program decided nothing on its own. Several of the cancers it found were small and caught early [1].
In February 2025 the UK's health department announced a much bigger trial, called EDITH, to test this kind of program across the national screening service: around 700,000 women at about thirty screening centres, with roughly £11 million of public research funding. The design keeps a person in the loop: the trial tests the program doing the job of one of the two doctors, never both [2].
Where it helped — because of where it sat
Look at where the program was placed. It checked the scans after two people had, not instead of them, and all it could do was ask a question — look at this one again — to a doctor who still owned the answer [1]. A second pair of eyes in that position can only add something — like a second cashier re-adding your bill before you pay: they can catch a mistake, but they cannot charge you more. The worst it can do is send a doctor back to a scan they had already cleared, which costs a minute. That is also why the big trial was set up the way it was. The program takes one of the two chairs, a person keeps the other, and the whole arrangement is measured against two people before anything changes for patients [2].
Where it burned — because of where it sat
Now the same kind of tool, in the opposite chair. Sepsis is the body's runaway reaction to an infection; it can kill within hours if nobody spots it, so hospitals want an early warning. Epic, a company whose software runs many hospitals' records, built a program to raise that alarm. It was in use at hundreds of US hospitals when researchers at the University of Michigan's hospital checked it against their own patients — nearly forty thousand hospital stays. The program missed about two out of every three sepsis cases it was meant to catch. Of the cases it did catch, most had already been spotted and treated by the staff; it gave a genuinely new warning only a small fraction of the time. At the same time it raised an alarm for almost one in five of all the patients in the hospital, so the nurses and doctors it was warning learned to look past it. It was like a car alarm that goes off every time a lorry drives past: after a week, nobody looks up. The company had reported a far higher accuracy, and hospitals had taken its word for it [3][4]. Same kind of program as the scan-checker. The difference was the chair: one sat next to people who still decided, the other sat in front of people who had no way to check it and every reason to stop looking.
The tell
Before you trust what an AI tool tells you at work, ask which chair it is sitting in. Is it a second pair of eyes — next to a person who still makes the call, and who can only be asked to look again — or is it the only pair of eyes, whose answer people act on straight away? If it is the only one, ask who outside the company that sells it has tested it doing that exact job, on people like yours.
Two systems can run the same kind of program and be nothing alike, and the difference is not in the program. The scan-checker could only ever add a second look, so when it was wrong it cost a doctor a minute. The sepsis alarm stood alone in front of tired staff. When it was wrong it cost their attention; when it missed, it cost a patient; and nobody at the hospital could measure any of that until researchers did the work themselves. So when a tool is offered to you — a flag on a scan, a fraud score, a review of your code, a summary you are about to forward — ask where it sits: is it the proofreader, or the author? A proofreader who misses a typo costs you a typo. An author you never read behind costs you the whole letter. A second pair of eyes beside your own judgment is cheap to be wrong. Something that replaces your judgment is expensive to be wrong, and the seller's own accuracy number is the one figure you cannot accept as the answer.
The check is a habit, and habits are trained. What to never hand off is about drawing the line between what an AI can draft and what only you should sign off — which is exactly the difference between a second pair of eyes and the only pair.
Sources
Every source below was opened and read. Last verified 27 September 2026.
- [1] AI helps doctors spot cancers missed in breast screening (NHS Grampian / Kheiron Mia evaluation) — BBC News, March 2024
- [2] AI breast cancer screening trial (EDITH) announced for the NHS — Department of Health and Social Care, UK Government, February 2025
- [3] External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients — JAMA Internal Medicine, 21 June 2021
- [4] Widely used sepsis prediction tool is less accurate than claimed, study finds — Michigan Medicine, June 2021