← All updates
Meta-analysis

Surgery · 3 h ago

Surgical AI Model Discrimination Varies by Dataset and Prediction Task

A review and quantitative analysis of 75 studies covering 295 surgical prediction models found lower average discrimination for NSQIP-derived models than for models using other datasets. Differences varied by outcome, and the adjusted association was no longer evident with precision weighting.

Researchers compared adult surgical artificial intelligence prediction models trained on the American College of Surgeons National Surgical Quality Improvement Program (NSQIP) with models using other datasets. Structured searches of PubMed, Scopus, and EMBASE covered publications from October 2015 through January 2026. The analysis included 75 studies contributing 295 models—145 NSQIP and 150 non-NSQIP—and assessed discrimination using area under the receiver-operating characteristic curve (AUROC).

Mean AUROC was 0.71±0.09 for NSQIP models versus 0.80±0.10 for non-NSQIP models (p<0.001). After adjustment for study and model characteristics, NSQIP use remained associated with lower AUROC (β=−0.056; 95% CI, −0.105 to −0.007; p=0.026). The largest difference occurred in infection prediction, with median AUROCs of 0.68 versus 0.98 (p<0.001). Mortality prediction showed no significant difference: 0.79 versus 0.80 (p=0.94).

However, precision weighting removed the adjusted association (β=0.008; 95% CI, −0.093 to 0.109; p=0.876), limiting confidence in an overall advantage for either dataset category. These cross-study comparisons assess discrimination, not improved clinical outcomes. The findings support matching training data to the intended prediction task rather than assuming that database size or standardization guarantees better performance.

AI summary · Not yet editor-reviewed

Is this summary clinically accurate?

Help fellow clinicians: your rating sends inaccurate summaries straight to our editors.

Sign in to rate this summary →

Source

Journal of the American College of Surgeons: Matching Surgical Datasets to Prediction Tasks: NSQIP and Non-NSQIP Artificial Intelligence Models Across Specialties ↗

This is an automated AI-condensed summary that has not yet been reviewed by an editor. Always consult the full item at the original source.