Glass-Box Governance: Evaluating Agentic AI in Public Health
More details
Hide details
1
Clinical AI, Wimmy (Pty) Ltd, Cape Town, South Africa
Popul. Med. 2026;8(Supplement Supplement 1):A930
ABSTRACT
INTRODUCTION:
As AI systems in health shift from single-turn chatbots to tool-using, stateful "agents," evaluation based only on final outputs is necessary but insufficient. We examine emerging approaches to auditing agent execution traces (actions, tool calls, retrieved evidence, memory updates, and human overrides) to support accountability in clinical and public health workflows.¹,²
METHODS:
We conducted a multivocal literature review of evaluation and governance approaches for LLM-based agents in healthcare and public health, synthesizing peer-reviewed studies, benchmarks, and regulator or standards guidance. We coded sources for evaluation targets (pipeline vs artefacts), assessment level (component vs system), and monitoring recommendations across the perception-reasoning-action loop.¹⁻⁴
RESULTS:
Across sources, accountability depends on system-level evaluation of end-to-end task trajectories rather than single responses.²,³ Trace instrumentation enables failure mode analysis and post-incident review, supporting organisational governance.¹,⁴,⁵ Verifier or judge patterns can improve reliability when verifiers apply explicit rubrics or tool-based checks.³,⁶ Evaluation should extend beyond static benchmarks to operational monitoring, including task completion rate and human intervention frequency, stratified by severity and drift over time.¹,⁴,⁵
CONCLUSIONS:
In the agentic era, public health organisations should adopt glass-box governance: trace-based observability, predefined acceptance criteria, continuous monitoring, and adaptive evaluation cycles that record failures, mitigations, and corrective actions to maintain alignment with clinical and population-health goals.¹,⁴,⁵