Noesa
← The Casebook

Thailand · 2018–2020

90% accurate in the lab. Then it met the lighting in a real clinic.

A screening model that matched specialists on clean data rejected more than a fifth of the photographs nurses actually took.

Published 15 August 2026

What happened

Diabetic retinopathy is an eye disease caused by diabetes. It leads to blindness, and it is treatable if caught early. That makes screening valuable: there are far more patients who need their eyes checked than eye specialists to check them. Google built an AI system to read photographs of the back of the eye and spot the disease. In the lab it was more than 90% accurate — what the team called "human specialist level" [2].

Researchers then did something unusual. They studied it in the field, across eleven clinics in Thailand, watching and interviewing the nurses who actually used it. Emma Beede and colleagues published the study at CHI 2020, a major conference on how people use technology [1][2].

In those clinics the system rejected more than a fifth of the photographs the nurses took. It had been trained on high-quality images and refused anything below its quality bar. The main reason was lighting: real examination rooms are lit however they happen to be lit [2].

The rest of the environment got in the way too. Slow internet made uploading images take a long time, frustrating staff and patients waiting for results. Nurses who could see that an eye looked healthy had their photograph rejected anyway, and patients were sent to another clinic on another day [2].

The published paper frames the finding not as a fault in the model but as a mismatch: tensions between the model's demands for data quality and the quality of data that a real, under-resourced clinic can produce [1].

Where it helped

When conditions let it work, it worked: an instant result in the clinic instead of a wait for a specialist, in a screening programme that could never hire enough eye doctors to see everyone. And the study itself is the helped story. A team published, in detail, the ways their own celebrated system fell short in the field. That is how a gap like this gets closed rather than repeated, and it is rarer than it should be [1][2].

Where it burned

Nothing here was a mistake in the model. The accuracy figure was honest, the training was sound, and the system did exactly what it was built to do — refuse to guess on a poor image, which is the responsible thing. It simply had never been asked to work in a room with the wrong lamp. The gap was not between the model and the truth. It was between the data it was tested on and the data it would actually be given, and no amount of extra accuracy in the lab would have shown it [1][2].

The tell

Before trusting a performance number, ask what it was measured on — and whether that looks like what you are actually going to feed it.

Every accuracy claim is a claim about a particular set of inputs. But the number gets quoted without them, so "90% accurate" travels as if it were a fact about the tool rather than about the test. It is like a car's advertised fuel economy: measured on a smooth test track, not on your hill in traffic. Your inputs are messier than the test's. Your documents are scanned crooked, your data has gaps, your users type badly. So the question is never just "how good is it?" It is "how good is it on things like mine?" — and the honest answer is often that nobody has measured that yet.

Share this case

The image has the link printed on it, so it still leads back here.

The check is a habit, and habits are trained. Product decisions, understood is about the distance between a thing that demos well and a thing that works where it is used — and how to tell which one you are looking at.

Sources

Every source below was opened and read. Last verified 15 August 2026.

  1. [1] A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic Retinopathy — Beede et al., Proceedings of CHI 2020 (Google Research), 2020
  2. [2] Google's medical AI was super accurate in a lab. Real life was a different story. — Will Douglas Heaven, MIT Technology Review, 27 April 2020