In March 2018, three colleagues and I published a short editorial in Bone & Joint Research. It was 2,000 words. It aimed to explain to a musculoskeletal readership what deep learning was, why it had recently begun to work and what it was likely to do to medicine. We closed with the following, which I still stand by. Medicine is much more than a diagnosis and treatment algorithm, and machine learning looked set to transform and complement how we deliver care.
Seven years on, I want to grade it. Some of what we predicted was right in the specifics. Some was directionally right, but for reasons we did not understand at the time. And two things happened that we could not have foreseen and that changed the field more than anything we described.
The point of the exercise is not vindication. Two of the four of us on that paper are still knee-deep in the field, one of my co-authors (David Golan) as co-founder and CTO of Viz.ai and one (me) as a surgeon and co-founder of Viz.ai. I have spent the decade watching medical AI go from a paper you had to explain to your colleagues into a purchase order that hospitals now consider at budget meetings. I want to know which parts of my old thinking survived contact with what actually happened.
What we got right
The first prediction was that regulatory bodies would move from permitting AI as clinical decision support to permitting it as diagnostic decision-making, and that clinicians should prepare for this rather than resist it. In 2018, the FDA had cleared a handful of AI/ML-enabled medical devices. By the end of 2024, the running list had passed a thousand, most of them in radiology. Every one of those clearances turned on the same argument our 2018 editorial gestured at: that a machine can, for a narrow well-defined task, meet or exceed a human specialist. The pace of clearance was faster than any of us expected. The direction was exactly right.
The second was that segmentation, image reconstruction and quantitative image analysis would move from research to routine practice. Every modern PACS now ships with automated organ and lesion segmentation. Cartilage volumetrics on knee MRI scans, once a research task, sits in the workflow. Automated bone age estimation on hand radiographs is a tool my residents can use without thinking. The paper predicted this specifically for musculoskeletal imaging. It happened.
The third was the point about geographically underserved populations. The DDH reference in that paper is my own group's work on automating Graf's method with deep convolutional networks. The case we made was that ultrasound expertise is scarce, that DDH is a screen-detected disease and that a model good enough to triage would let unskilled users image children well and refer only those who needed expert eyes. The full deployment we imagined has not happened yet. But the modality has moved. Automated musculoskeletal ultrasound is now an available tool in screening programmes. The clinical case was right. The economics still lags.
The fourth thing we got right is one I am quietly proud of. We framed the paper around narrow AI, and we said explicitly that general AI, machines that replicate human thought, would remain in the realm of science fiction. In 2018, that was the professional consensus. We were wrong about the specifics of what came next, and I will get to that. But the framing held. The AI that has actually changed medicine is still, structurally, narrow. It reads a scan. It flags a stroke. It transcribes a consultation. Even when the underlying model is general, its deployment is narrow, tuned end to end on one problem or prompted to behave as if it were. Nobody has an AI in a hospital that is trying to replicate the full interior workings of a doctor. Yet.
What we got directionally right but for the wrong reasons
We predicted that machine learning would transform radiology first. It has, and every measure of AI deployment in medicine shows radiology in the lead. But the reason we gave was that radiology had the imaging volumes and the pattern-recognition tasks. That is true and incomplete. The real reason radiology led is that radiology already had a fully digitised workflow, standardised outputs in DICOM, a clear billing model and consultants who reported directly to referrers rather than to patients.
Pathology has similar volumes and pattern tasks but is only now catching up, because the workflow was largely analogue until slide scanning became affordable. Ophthalmology moved fast for the same reason: fundus photography is DICOM-adjacent and reads out to specialists. The lesson, which I did not internalise then and which I have since, is that clinical AI adoption is not gated by clinical need or by model performance. It is gated by workflow readiness. That single insight would have saved me and quite a few investors a lot of time.
We also cited IBM Watson Health, uncritically, as the exemplar of the coming order. IBM sold Watson Health to a private equity firm in January 2022 for around a billion dollars, having invested tens of billions and generated most of the industry's early scepticism about medical AI in the process. The failure came from betting on the wrong shape of product: a monolithic decision-support system layered over existing electronic health records, with no clear who-pays economics and no measurable outcome it could defend. It also came from marketing the product as an oracle rather than a tool. The Watson autopsy is now the standard first-week reading for any clinical AI company I advise. I do not regret citing them in 2018. In 2018, they were the loudest signal in this space. I regret reading the noise as evidence.
What we could not have predicted
Two things happened after our paper that reshaped the field.
The first is the transformer. In June 2017, a group of Google researchers published Attention Is All You Need. It described an architecture called the transformer, and it was going to eat every previous approach in machine learning within five years. In 2018, if you had asked me what would drive the next decade of medical AI, I would have told you convolutional neural networks and better labelled datasets. Those were the tools that were working. We did not know a paradigm shift had just landed. GPT-3 arrived in 2020, GPT-4 in 2023, and by 2024, the interesting question in medical AI was no longer what task-specific model to build but which foundation model to fine-tune, prompt or ground against your data.
The general-versus-narrow dichotomy we opened our 2018 paper with turned out to be too clean. What emerged sits between the two: models trained on everything, prompted to behave narrowly, with performance that on some tasks matches the specialists we spent a decade trying to build task-specific models to catch. If I were writing that paragraph today, general AI would not be filed under science fiction. It would be filed under a live regulatory question.
The second thing we could not see is what I have come to think of as the evidence gap. In 2018, the assumption in the room was that if you could get a model to perform, deployment would follow. Show the accuracy and the ROC curve, and the hospital will buy it. That has turned out to be exactly wrong.
Between 2018 and 2026, the accuracy of task-specific models became a commodity. Any competent team, using open weights and off-the-shelf tooling, can now build a respectable model for most defined clinical tasks in a matter of weeks. What is scarce, and what actually determines whether a company sells anything, is proof that the model changes an outcome the payer already cares about, together with integration deep enough that the output is put in front of the right person at the right time.
Viz.ai's stroke platform did not become a success because its LVO detector was uniquely accurate. Plenty of groups had models that could find a blocked vessel. It became a success because it changed door-to-treatment time and was the first AI product to earn a CMS New Technology Add-on Payment, a reimbursement pathway in the US that most of the industry did not even know existed in 2018.
Our 2018 paper mentioned diagnostic accuracy and cost reduction as the twin objectives of clinical AI. Neither on its own has proven sufficient. The industry has spent years painfully rediscovering what a good health economist could have told us in the first paragraph: what a hospital buys is a change in an outcome that flows through its own budget. Everything else is a nice-to-have demo.
What survives and what I would rewrite
The single sentence from the 2018 editorial that I would still write today, unedited, is the closing one: medicine is much more than a diagnosis and treatment algorithm, and machine learning looks set to transform and complement how we deliver care. Every word of that has held true. The transformation has been faster than we expected in imaging and slower than we expected everywhere else. AI as a replacement for a clinician is not a great sell. AI as an amplifier of a clinician is the current trajectory.
The sentence I would rewrite is the one that placed general AI in the realm of science fiction. It sits in the room now, in the browser tab I have open next to me while I write this. It has not replaced any of my judgement in the operating theatre yet, but it can read my scans before I do and draft my letters. It can start to argue its own case for what the diagnosis is. What I would say today is that the general-narrow dichotomy has collapsed, and the useful question is whether a given deployment has been validated against a specific outcome in a specific population, whatever the underlying architecture happens to be.
The paper we wrote in 2018 was a first sketch of a field that was just becoming visible to clinicians. It got the trajectory broadly right and the mechanisms partly wrong. The single missing sentence, which I will write in an essay called The Evidence Constraint, is this: model performance is no longer where the value is created in medical AI. It has not been for some years. The value is created downstream, in integration depth and outcome evidence. The next decade of clinical AI is going to be won by whoever solves those two problems. Everyone else is going to spend it explaining why their models were accurate.
That is the argument I will continue to make. We were right that this technology would transform medicine. We were wrong about where in medicine it would matter, and about what would separate the winners from the also-rans. What is not scarce is a model that can pass a benchmark. What is scarce, still, is proof that a model changes the number a payer cares about, delivered in front of the clinician at the exact moment they can act on it.