Whose Data, Whose Benefit? Equity Implications of AI-Based Surgical Risk Prediction Models
 
More details
Hide details
1
School of Medicine, Nelson Mandela University, Gqeberha, South Africa
 
2
Department of Health, Livingstone Tertiary Hospital, Gqeberha, South Africa
 
3
Department of Health, Hermanus Hospital, Hermanus, South Africa
 
4
School of Medicine, University of Pretoria, Pretoria, South Africa
 
5
Department of Computing, Mathematical, and Statistical Sciences, School of Science,, University of Namibia, Windhoek, Namibia
 
6
Faculty of Health Sciences, Nelson Mandela University, Gqeberha, South Africa
 
 
Popul. Med. 2026;8(Supplement Supplement 1):
 
ABSTRACT
BACKGROUND:
Accurate surgical risk prediction is central to safe perioperative care and increasingly supported by artificial intelligence (AI). While AI-based models show strong performance, their development and validation are not value-neutral. From a public health equity perspective, reliance on evidence predominantly from high-income countries (HICs) raises concerns about relevance and benefit for low- and middle-income countries (LMICs), where unmet surgical need is greatest.

METHODS:
We conducted a systematic review and meta-analysis of AI-based surgical risk prediction models. PubMed, IEEE Xplore, Scopus, and Web of Science were searched from Jan 1, 2019, to June 30, 2025, for studies predicting perioperative complications, using MeSH terms and free-text keywords related to AI, machine learning, deep learning, logistic regression, random forest, support vector machines, gradient boosting, and surgical risk outcomes. Model discrimination was pooled using random-effects meta-analysis.

RESULTS:
Of 8969 citations, 48 studies from 25 countries, comprising 6,959,207 patients were included, with six multi-country studies. Evidence was heavily skewed toward high-income-countries (HICs) (n=22), with limited representation from upper-middle-income (UMIC) (n=2) and LMICs (n=1). Overall, AI models demonstrated high discriminative performance (pooled AUC 0.894, 95% CI 0.881–0.907), with similar performance across modelling approaches. Gradient boosting achieved the highest discrimination (AUC 0.918, 95% CI 0.893–0.944; k=7). Subgroup analyses showed comparable performance for machine learning (AUC 0.888), deep learning (0.882), random forest (0.892), logistic regression (0.890), and support vector machines (0.894). Reporting of sociodemographic variables, health-system context, and external validation in LMICs was sparse.

CONCLUSIONS:
AI-based surgical risk prediction models perform well, but their evidence base is largely drawn from well-resourced health systems. This imbalance raises equity concerns regarding generalizability and fairness. Without intentional inclusion of diverse populations and health-system contexts, AI innovations may disproportionately benefit advantaged groups. Equitable application requires LMIC data, context-specific validation, and transparent equity-focused reporting.
eISSN:2654-1459
Journals System - logo
Scroll to top