Radiology AI is bottlenecked by data that does not exist, not by the models themselves
The aggregated training dataset required to build a radiology AI that covers the full range of imaging types has never been assembled. Until it is, developers are confined to narrow subdomains, and the economic case for the technology is weaker than its proponents acknowledge.
Radiology AI has a coverage problem, and Eric Vishria names the cause directly. The aggregated training dataset that would allow a model to perform across the full spectrum of imaging types does not exist. Without it, developers are forced into narrow subdomains, building tools that handle one variety of scan well while leaving the rest untouched.
That narrowness matters because of what radiologists actually do each day. Vishria puts the range at 20 to 40 different scan types per shift. An AI that handles only one of those is, as he frames it, only marginally useful. The tool may perform well inside its lane, but a radiologist cannot hand off a meaningful portion of their workload to something that covers so little of it.
The missing-data problem is not specific to radiology. Stacey Stephens, describing the parallel challenge in robotics, notes that language models benefit from training on the entire body of data available on the internet, while that kind of at-scale data simply does not exist for robot control. The structural condition is the same: where real-world physical or clinical data has never been systematically collected and labeled, there is no shortcut to assembling it. Researchers must generate it from scratch, which is slow and expensive.
The big hurdle and the big thing that they articulated was wait a minute all of the aggregated training data set doesn't exist anywhere Eric Vishria
The data constraint also helps explain why coding has been such a visible early success for AI while most other domains have lagged. David George, an investor, makes the distinction plainly: coding is perfectly documented, verifiable, and simulatable. Most business tasks, and most clinical tasks, share none of those three attributes. Radiology sits squarely in that harder category. Imaging data is governed by strict privacy regulations, fragmented across institutions, and rarely labeled in ways that transfer cleanly between sites or scanner types. What makes coding tractable is precisely what makes radiology hard.
Even setting aside the coverage gap, the economic argument for radiology AI carries a structural weakness. Harry Stebbings observes that automating the bulk of radiology tasks may not translate into workforce reduction at all. The residual work, the cases too complex or ambiguous for the automated system to handle, can be sufficient on its own to justify keeping the same number of radiologists. If imaging volume keeps rising, even a system that handles 95 percent of reads can leave the remaining five percent as a full-time job.
Taken together, these constraints describe a technology that is real and in some contexts genuinely useful, but that is further from transforming the field than the ambient optimism around medical AI suggests. The path to broad radiology coverage runs through a data aggregation problem that no single company can solve unilaterally, and the path to economic disruption runs through a residual-work problem that grows alongside demand. Neither is intractable in principle. But neither has a near-term solution that the evidence currently supports.
The analogy to robotics is worth holding onto. In both cases, the limiting factor is not the model architecture or the compute budget. It is the prior question of whether the training signal exists at a scale that makes generalization possible. Answering that question requires institutional coordination, regulatory accommodation, and years of labeling work. Until that groundwork is laid, radiology AI will remain a collection of narrow, high-performing tools in search of a workflow broad enough to justify their adoption at scale.