Logistic Regression vs. Machine Learning: Simulation Insights into Predictive Accuracy for Occupational Health Outcomes
 
More details
Hide details
1
Graduate Program in Collective/Public Health, Botucatu Medical School - São Paulo State University (UNESP), Botucatu, Brazil
 
2
Department of Public Health, Botucatu Medical School - São Paulo State University (UNESP), Botucatu, Brazil
 
3
Graduate Program in Nursing Academic Master’s and Doctoral Programs, Botucatu Medical School - São Paulo State University (UNESP), Botucatu, Brazil
 
4
Department of Sociology, Social Work and Public Health - Faculty of Labour Sciences, University of Huelva, Huelva, Spain
 
5
Safety and Health Postgraduate Programme, Universidad Espíritu Santo,, Guayaquil, Ecuador
 
6
Department of Preventive Medicine and Public Health, Sevilla Medical School - University of Sevilla, Sevilla, Spain
 
 
Popul. Med. 2026;8(Supplement Supplement 1):
 
ABSTRACT
ABSTRACT:
This study aimed to compare the performance of logistic regression (LR) and targeted maximum likelihood estimation (ML) in predicting return-to-work outcomes using simulated datasets derived from a historical cohort of 738 Brazilian public university employees. To emulate real-world epidemiological conditions, a synthetic population of 10,000 cases was generated based on empirical distributions of two exposures—changes in ICD-10 diagnosis chapter (81%) and occurrence of a mental health diagnosis (44%)—alongside multiple sociodemographic and occupational covariates. From this dataset, 100 simple random samples were drawn in four sizes (n=150, 750, 1,500, and 5,000) to assess estimator behavior across varying levels of statistical power. Analyses demonstrated marked differences between LR and ML. For the exposure “modification of ICD-10 chapter” (Model 1), the population-adjusted OR was 0.28. LR consistently produced estimates close to this value across all sample sizes (0.28–0.29), with RMSE decreasing sharply as sample size increased (from 0.24 at n=150 to 0.02 at n=5,000). ML produced systematically higher ORs (0.44–0.46) and larger RMSEs (0.27 to 0.16), with the true OR included in ML confidence intervals only at the smallest sample size. For mental health diagnoses (Model 2), the population-adjusted OR was 27.35. LR estimates converged toward this value as sample size increased (60.9 at n=150 to 27.2 at n=5,000), whereas ML yielded more stable but substantially conservative ORs (15.3–17.4). RMSE values again favored LR, decreasing from 96.8 (n=150) to 1.97 (n=5,000), while ML RMSE remained comparatively high and stable (12.0–13.0). These findings demonstrate that under simulated epidemiological scenarios with predefined causal structure, LR often provides more accurate and precise estimates than ML—particularly with adequate sample sizes. Contrary to assumptions of algorithmic superiority, ML did not outperform LR in estimating exposure–outcome associations, underscoring the continued relevance of traditional regression approaches in occupational health research.
eISSN:2654-1459
Journals System - logo
Scroll to top