A Practical Framework for Data Science in Epidemiological Surveillance
 
More details
Hide details
1
Department of Health Administration and Planning, Oswaldo Cruz Foundation (Fiocruz), Rio de Janeiro, Brazil
 
2
Office of Fiocruz Strategy for the Agenda 2030, Oswaldo Cruz Foundation (Fiocruz), Rio de Janeiro, Brazil
 
3
Computing Department, São Paulo State University, São Paulo, Brazil
 
4
Department of Epidemiology and Quantitative Methods, Oswaldo Cruz Foundation (Fiocruz), Rio de Janeiro, Brazil
 
 
Popul. Med. 2026;8(Supplement Supplement 1):A891
 
ABSTRACT
INTRODUCTION:
The rapid expansion of digital health data and analytical technologies has positioned data science as a key enabler of modern epidemiological surveillance. Machine learning, big data analytics, and artificial intelligence support early outbreak detection, risk stratification, and evidence-informed public health decision-making.¹–³ Despite this potential, many initiatives still rely on ad hoc or poorly documented processes, limiting reproducibility, transparency, and scalability.⁴,⁵ This study proposes a structured methodological workflow to support the systematic application of data science in epidemiological surveillance.

METHODS:
This study was informed by a scoping review conducted according to the “Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews” (PRISMA-ScR) guidelines.⁶ Peer-reviewed literature addressing data science applications in epidemiology and public health surveillance was identified across major international databases. Included studies were analyzed to identify recurring phases, activities, roles, technical requirements, deliverables, and governance practices. Based on this synthesis, an integrative workflow model was developed, drawing on established data science methodologies and public health principles.⁴,⁷,⁸

RESULTS:
The proposed workflow comprises six interconnected stages: i) data collection and acquisition, ii) data preprocessing, iii) data provisioning, iv) data modeling and analysis, v) use and dissemination of results, and vi) data governance. Each stage is defined by specific objectives, core activities, and expected outputs. Data governance is treated as a transversal component, addressing data quality, privacy, security, interoperability, traceability, and reproducibility throughout the workflow.³,⁸,⁹ The model is adaptable to different epidemiological contexts, data sources, and analytical goals, including routine surveillance and emergency response scenarios such as pandemics.¹⁰–¹²

CONCLUSIONS:
A structured workflow enhances rigor, transparency, and reproducibility in data science projects for epidemiological surveillance. The proposed model offers a practical and internationally applicable reference for researchers, public health professionals, and institutions seeking to operationalize data science in a systematic, ethical, and decision-oriented manner, strengthening global public health surveillance capacities.¹³
eISSN:2654-1459
Journals System - logo
Scroll to top