groundy
Culture & Society

AI Diagnostics in 2026: Where Machines Now Outperform Radiologists

AI tools beat radiologists on narrow imaging tasks like lung nodule detection, yet only about 2% of U.S. radiology practices use them. The evidence, gaps, and barriers.

Published Updated 10 references
On this page7 sections

AI diagnostic tools outperform radiologists on specific, narrow imaging tasks. On lung nodule segmentation from CT scans, a deep learning system achieved an AUROC of 94.4% and outperformed six radiologists on the same task (Redefining Radiology. PMC, 2023); on a 936-case diagnostic challenge, a multimodal research model outscored physician respondents 61% to 49% (AI in Radiology: 2025 Trends, FDA Approvals & Adoption. IntuitionLabs, 2025). The performance evidence is peer-reviewed. Deployment is the lagging half. A 2024 European Society of Radiology survey of 572 respondents (IntuitionLabs, 2025) found 48% actively using AI tools, up from 20% in 2018, while an Associated Press estimate from the same period put AI reading-tool integration at only about 2% of American radiology practices (IntuitionLabs, 2025).

The Performance Gap Is Real and Specific

Precision matters here. AI does not uniformly outperform radiologists across all imaging tasks. It wins in particular, well-defined contexts: high-volume screening, pattern recognition on standardized imaging, and detection of anomalies that reward tireless, consistent attention. Understanding where the performance gap exists is the prerequisite to any serious discussion of adoption.

The strongest head-to-head results come from narrow detection tasks. For lung nodule segmentation on CT, AI systems recorded an AUROC of 94.4%, outperforming a comparison cohort of six radiologists on the same task (PMC, 2023). On musculoskeletal radiographs, a calibrated ensemble of deep learning models beat expert radiologists in three of seven upper-extremity regions, with an AUC of 0.93 (PMC, 2023). Three of seven is the honest headline: the ensemble outperformed expert readers in three regions and not in the other four, and the review reports that split without explaining what separated the winning regions.

Mammography results are more nuanced. The same review summarizes AI models distinguishing benign from malignant lesions with performance “comparable to human radiologists,” and notes AI has shown promise in “equalling, and in some cases surpassing” radiologists in breast screening (PMC, 2023). The concrete, quantified gain is in false positives: in a retrospective comparison, an AI-based CAD system reduced false-positive marks per image by 69% against traditional CAD software (PMC, 2023), cutting microcalcification marks by 83% and mass marks by 56%, with reading time down an estimated 17% (PMC, 2023). Neither side is categorically better; the tradeoff is between catching more cancers and generating fewer unnecessary callbacks.

General-purpose models show the same pattern from the other direction. The multimodal GPT-4V model achieved 61% accuracy on a 936-case diagnostic challenge, outscoring physician respondents at 49% (IntuitionLabs, 2025), and one report credited it with 85% accuracy detecting multiple sclerosis progression on brain MRI (IntuitionLabs, 2025). Those are research results, not products: as of late 2025, no large language model is FDA-authorized for clinical imaging use (IntuitionLabs, 2025).

Where AI Has the Clearest Edge

The following table summarizes the current evidence by task:

TaskAI PerformanceComparatorFinding
Lung nodule CT segmentationAUROC 94.4% (PMC, 2023)6-radiologist cohortAI outperforms cohort
MSK radiograph abnormality detectionAUC 0.93 (PMC, 2023)Expert radiologistsAI leads in 3 of 7 upper-extremity regions
Mammography false positivesMarks per image down 69% (PMC, 2023)Traditional CAD softwareFewer false positives; reading time cut about 17%
Chest X-ray triageReport delivery cut from 11.2 days to 2.7 days (PMC, 2023)Baseline workflowAutomated triage speeds interpretation
Stroke alert workflowTreatment about 66 minutes sooner (IntuitionLabs, 2025)Pre-AI workflowViz.ai deployed in 1,600+ hospitals
General diagnostic challengeGPT-4V 61% accuracy (IntuitionLabs, 2025)Physicians at 49%Research model, not an authorized product

One operational result sits outside the diagnostic columns because it is a workflow outcome: Viz.ai’s stroke-alert platform, deployed in more than 1,600 hospitals, is associated with stroke treatment starting about 66 minutes sooner (IntuitionLabs, 2025). European regulators have also begun allowing fully automated normal-versus-abnormal chest X-ray triage, though typically inside workflows that keep a human reviewer in the loop (IntuitionLabs, 2025).

The pattern is not that AI is better. It is that AI is demonstrably stronger on narrow, high-volume, pattern-dense detection tasks where consistency matters. Every tool-based result in the table is a detection, triage, or workflow measure; the one open-ended row, the general diagnostic challenge, came from a research model with no FDA authorization as of late 2025, so it does not belong in the same column.

Digital Pathology: A Quieter Revolution

While radiology gets most of the headlines, digital pathology consolidated sharply in 2025. Paige trained a foundation model on more than one million pathology slides with Microsoft and released it as open-source; Tempus announced and completed its acquisition of Paige in August 2025, so Paige is no longer a standalone vendor (Top Imaging & Pathology AI Companies. IntuitionLabs, 2025). Quest Diagnostics completed its acquisition of PathAI’s CLIA lab and licensed its AISight digital-pathology platform in June 2024, and AISight received its CE mark in 2025 (IntuitionLabs, 2025). Proscia raised $50 million in 2025, part of $130 million total, counts 16 of the top 20 pharmaceutical companies among users, and processes roughly 22,000 patients per day on its Concentriq platform (IntuitionLabs, 2025).

The clinical evidence underneath those deals has the same shape as radiology’s. Reviews report that AI algorithms in tissue analysis have “markedly enhance[d] diagnostic accuracy and speed,” flagging subtle histopathologic attributes often overlooked by the human eye (PMC, 2023).

Deployment lags here too. Surveys cited in market analyses show only a few high-volume labs are fully digital (IntuitionLabs, 2025). And the regulatory path is not lighter than radiology’s for diagnosis: the FDA regulates AI intended to diagnose disease as a medical device, with pathology AI assisting tissue analysis and cancer detection among the authorized categories, while certain administrative and wellness software is excluded from device regulation under the 21st Century Cures Act (Akin Gump, 2025).

Why Hospitals Aren’t Moving

The adoption gap is regional as much as institutional. The contrast between 48% usage among surveyed European radiologists and roughly 2% of U.S. practices is not a difference in evidence; it is a difference in environment (IntuitionLabs, 2025). The same European survey found about 80% of respondents did not understand how AI medical devices are approved or monitored (IntuitionLabs, 2025), which points to a training gap as much as a skepticism gap. Hospital-level numbers sit above practice-level ones: a 2024 study published in JAMA Radiology (IntuitionLabs, 2025) found roughly 20% of major U.S. hospitals had piloted at least one AI radiology tool, with 10% using one clinically. Large systems move first; small practices do not. Even where activity is broad, clinical use stays narrow: imaging and radiology was the most widely deployed clinical AI use case in a survey of U.S. health systems (Adoption of artificial intelligence in healthcare. PMC, 2025), with 90% of responding organizations reporting at least partial deployment, but successes with diagnostic use cases were limited.

The reasons U.S. deployment lags are structural, not informational:

Tool maturity concerns: In a survey of U.S. health systems with 43 responding organizations (PMC, 2025), 77% cited lack of AI tool maturity as the biggest or second-biggest barrier to deployment, ahead of financial concerns at 47% and regulatory or compliance uncertainty at 40%. The sample is small, so read the percentages as directional; the ordering is the point. Many approved tools were validated on curated datasets that do not match the messiness of real-world imaging: varying equipment manufacturers, patient demographics, and scanner settings.

Reimbursement gaps, narrowing in places: Medicare has no standard method for covering and paying for AI-enabled services; its coverage policies are item- or service-specific (Akin Gump, 2025). Narrow slices are opening: the 2026 Hospital Outpatient Prospective Payment System final rule established national payment for AI-assisted cardiac analysis, and the AMA’s CPT Editorial Panel has created Category I codes for AI-assisted retinal imaging analysis and cardiac imaging interpretation (Akin Gump, 2025). Where a billing code exists, adoption has a financial path. Where it does not, early adopters still carry the cost and the coverage risk.

Workflow integration: AI tools require PACS integration, staff training, and workflow redesign, and radiologist acceptance remains the gating factor (PMC, 2023). A system cleared by the FDA still requires that integration work before it touches a patient.

Regulation still being written: The landscape is one of ongoing legislative debate, FDA oversight, and shifting CMS payment policy rather than settled rules (Akin Gump, 2025). The EU’s AI Act, effective January 2026, classifies medical AI as high-risk and requires documented bias checks and human oversight (IntuitionLabs, 2025). In the U.S., an ONC final rule issued in December 2023 imposes transparency and risk-management requirements on predictive decision-support interventions (Akin Gump, 2025). A tool compliant in one jurisdiction can face different requirements in another.

The Liability Labyrinth

The legal framework for AI diagnostic errors has not caught up to the clinical reality. A systematic review of the medical liability literature on diagnostic algorithms concluded that the ethical and legal concerns raised by AI still have “no unanimous response” (Defining medical liability when artificial intelligence is applied on diagnostic algorithms. PMC, 2023). In most jurisdictions the radiologist signs the report and is legally responsible, and error attribution when an algorithm errs remains a legal grey area (IntuitionLabs, 2025).

Research through Brown University’s Alpert Medical School, published in NEJM AI in 2025, quantified what the study’s authors call the “AI penalty.” More than 1,300 participants acted as jurors on two hypothetical malpractice cases (Brown University Alpert Medical School, July 2025). In the missed brain-bleed scenario (Brown University, July 2025), participants sided with the plaintiff more than 56% of the time when no AI was involved, and half the time when both the AI and the radiologist missed the finding. When the AI flagged the bleed and the radiologist did not, plaintiff-side verdicts rose to nearly three-quarters (Brown University, July 2025). The AI becomes a witness against the clinician, not a shared liability partner.

Liability exposure fans out across three parties: the clinician, who retains final diagnostic authority regardless of AI output; the AI vendor, if the algorithm contains documented flaws; and the hospital, if the tool was not properly maintained or validated for the deployment context. The liability literature surveys these allocation problems, and one reviewed author argues a common enterprise strict liability approach “would create strong incentives for the relevant actors to take care” (PMC, 2023). This ambiguity discourages institutional deployment. Risk management departments face genuine legal uncertainty, and absent clear precedent, caution prevails.

The FDA’s Expanding Authorization Registry

The regulatory picture gives context to the market’s trajectory. The FDA’s March 2026 list update, covering authorizations through December 2025 (FDA Updates AI List with New Clearances. The Imaging Wire, March 2026), shows 1,451 AI-enabled medical devices authorized since tracking began in 1995, with 1,104 of them (76%) radiology tools. In Q4 2025 alone, the FDA cleared 72 AI-enabled devices, 55 of them radiology products (The Imaging Wire, March 2026). Radiology took 75% of AI authorizations across 2025 (The Imaging Wire, March 2026), against 73% in 2024 and 80% in 2023, so its lead is durable even as its share sits below the 2023 peak.

The leading cleared-device holders by volume, per the FDA list (The Imaging Wire, March 2026): GE HealthCare at 120 radiology AI authorizations, Siemens Healthineers at 89, Philips at 50, Canon at 45, United Imaging at 38, Aidoc at 31, and DeepHealth at 28. The list counts imaging hardware with embedded AI, such as mobile X-ray systems with detection algorithms, alongside standalone software. Most imaging AI reaches the U.S. market through the 510(k) pathway, cleared by demonstrating substantial equivalence to a predicate device (IntuitionLabs, 2025).

Philips’ SmartSpeed Precise, cleared in July 2025, shows the installed-base side of the list. It is a dual-AI deep learning reconstruction software upgrade for Philips’ existing 1.5T and 3.0T MRI systems, not new hardware; per Philips it delivers scans up to three times faster than SENSE protocols and images up to 80% sharper than SENSE and C-SENSE baselines across that installed base, though Philips notes the software is not yet CE marked (Philips Press Release, 2025).

Authorization volume is not deployment volume. Against 1,104 radiology authorizations, roughly 2% of U.S. practices (AI in Radiology: 2025 Trends. IntuitionLabs, 2025) and 10% of major U.S. hospitals (Top Imaging & Pathology AI Companies. IntuitionLabs, 2025) report clinical use of an AI radiology tool, so on those numbers most cleared tools have not reached routine deployment. The pipeline is full; the implementation infrastructure is not.

Why the Evidence Points to Augmentation, Not Replacement

Standalone accuracy is the wrong metric to optimize. Across surveys and observational studies, radiologist-AI combinations outperform either alone: combined human and AI reading slightly outperforms radiologists alone in cancer detection, with potentially lower false-negative rates (IntuitionLabs, 2025), and mammography CAD systems historically raised radiologist sensitivity by 5 to 10% on average (IntuitionLabs, 2025). The Grenoble trauma study’s split error pattern points the same way: two readers with different blind spots catch more than one (IntuitionLabs, 2025).

The field’s own consensus language is augmentation. Moody et al. (2025) conclude that “AI must amplify, not diminish, human capability to be effective,” and an AP News comparison that circulated widely likens AI to an aircraft autopilot: it can fly the plane, but a human pilot is still essential for safety, an analogy that captures radiologists’ common sentiment (IntuitionLabs, 2025). Reviews of AI integration similarly argue for a radiologist-AI feedback loop rather than autonomous reading (PMC, 2023).

This is the model most hospitals want to reach. The gap is between wanting it and having the legal, financial, and operational infrastructure to deploy it responsibly. Akin Gump’s practical recommendation runs in parallel: healthcare stakeholders should build AI governance frameworks covering validation, monitoring, and clinical oversight (Akin Gump, 2025).

For practitioners evaluating AI diagnostic tools, the questions that change a decision follow from the evidence above: generalizability (was this tool validated on your patient population and equipment, given the documented drop when imaging protocols differ?), workflow fit (does it integrate with your PACS and reporting structure?), and liability posture (does your institution track and disclose the tool’s error profile, given that disclosure measurably shifts juror outcomes?). Benchmark superiority on a curated dataset is a starting point, not a deployment decision.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy