Missing Data Imputation in the HAALSI Study: Comparative Analysis of Statistical, Machine Learning, and Deep Learning Techniques
More details
Hide details
1
MRC/Wits Rural Public Health and Health Transitions Research Unit (Agincourt), University of Witwatersrand, Johannesburg, South Africa
Popul. Med. 2026;8(Supplement Supplement 1):A3698
ABSTRACT
BACKGROUND:
Missing data is common in surveys and longitudinal studies, posing significant challenges for statistical analysis. Existing imputation approaches often struggle with datasets that include mixed variable types and complex missingness patterns. In addition, limited empirical evidence exists on how predictor selection influences imputation performance. This study compared statistical, machine learning (ML), and deep learning (DL) methods for missing data imputation and examined the role of variable importance in guiding predictor choice.
METHODS:
Using the HAALSI baseline dataset, a complete-case sample was derived, and missing values were artificially introduced to replicate the original missingness structure. Seven imputation methods were evaluated: statistical (Multiple Imputation by Chained Equations [MICE]), ML-based (Random Forest, XGBoost, Support Vector Machine [SVM]), and DL-based (CTGAN, CopulaGAN, Denoising Autoencoder [DAE]). Each method was repeated five times to assess stability. Performance for categorical variables was evaluated using accuracy, Cohen’s Kappa, sensitivity, and specificity, while numerical variables were evaluated using RMSE, MSE, and MAE. Variable importance analyses were conducted to identify key predictors for each imputed variable.
RESULTS:
SVM consistently outperformed other methods, achieving the highest accuracy (0.92 ± 0.038), sensitivity (0.99 ± 0.055), and Kappa (0.70 ± 0.041), indicating strong agreement and robust classification performance. XGBoost and DAE demonstrated higher specificity (0.77), suggesting improved discrimination of true negatives. Random Forest and MICE showed moderate performance. GAN-based approaches generally performed poorly. Variable importance analyses revealed that most variables required fewer than five predictors for imputation.
DISCUSSION:
Machine learning methods (i.e SVM) perform well under Missing at Random (MAR) conditions. These findings support growing evidence that kernel-based and neural network approaches are well suited to complex population health data. Future research should extend benchmarking across different missingness mechanisms and sample sizes to assess generalisability. Tailoring imputation strategies to data characteristics is essential for improving reproducibility and inferential validity in research.