Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity of supervised signals in multi-turn ophthalmic telephone triage, caused by high expert annotation costs and clinical privacy constraints. To overcome this challenge, the authors propose the Guideline-as-Oracle (GAO) framework, which translates the American Academy of Ophthalmology guidelines into 70 operational rules to serve as the sole instance-level supervision for 3,000 training dialogues—eliminating the need for manual annotation. By introducing eight novel rule-to-dialogue construction strategies, a rule-based zero-annotation dialogue generation method, a label repair mechanism, and fine-tuning a 9B-parameter language model, GAO-Triage achieves a significant improvement on a reference set of 201 cases: expert agreement rises from 61.7% to 74.1%, and recall for urgent cases increases dramatically from 9.5% to 69.0%. The system outperforms all general-purpose baselines without relying on advanced reasoning techniques.
📝 Abstract
Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted. We introduce Guideline-as-Oracle (GAO), which compiles American Academy of Ophthalmology guidance into a 70-row operational rule table and uses it as the sole source of instance-level supervision for 3,000 training dialogues, reserving human labeling for evaluation. Because converting rules into dialogues is itself a design problem, we catalog eight construction strategies, including cited-row tier assignment, one-fact boundary pairs, metadata-only repair, and label repair, and characterize the evidential status of each: labeling mechanism, null, confounded, or evaluated only as a package. Fine-tuning a 9B backbone on this corpus yields GAO-Triage, improving agreement with a 201-case operational reference from 61.7% to 74.1% (exact McNemar p=0.0046) and emergent-case recall from 9.5% to 69.0%; the gains persist across a second seed and patient simulator. None of the seven general-purpose systems we test dominates GAO-Triage on both metrics, and GAO-Triage requires no frontier model at inference time. Permuting label-dialogue assignments collapses the model to a constant-routine predictor, indicating that the signal lies in guideline-derived assignment rather than dialogue surface form. Label repair coincides with the disappearance of a late-training safety degradation.
Problem

Research questions and friction points this paper is trying to address.

zero-annotation
medical dialogue
telephone triage
supervision scaling
privacy-restricted
Innovation

Methods, ideas, or system contributions that make the work stand out.

Guideline-as-Oracle
zero-annotation training
rule-based dialogue generation
medical triage agent
label repair
🔎 Similar Papers
No similar papers found.
C
Chenyu Wang
Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences
Y
Yi Liu
ByteDance Inc.
B
Baoqing Li
Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences
M
Min Tu
Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences
D
Diping Song
Shanghai Artificial Intelligence Laboratory