Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses persistent challenges in long-form report generation by large language models—namely factual drift, internal inconsistencies, and lack of traceability. The authors propose a two-tiered agent architecture: a deterministic “Librarian” constructs a timestamped, trust-tiered knowledge base comprising evidence cards, authority-score ledgers, and claim graphs; a portable multi-agent “Writer” leverages this infrastructure to produce coherent, fully traceable reports by any deadline T. This framework achieves, for the first time, “align once, write anytime” capability, enabling strict temporal consistency replay and red-team feedback–driven auto-correction. Experiments demonstrate elimination of 6,845 cross-document contradictions across a corpus of over 550,000 evidence cards. A trust-prioritization strategy attains 100% accuracy on 22 gold-standard cases—substantially outperforming a 41% baseline—while exhibiting zero lookahead violations, 3.7× faster inference, and superior output quality relative to an all-Opus baseline.
📝 Abstract
Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing. A deterministic "librarian" ingests timestamped sources into a trust-tiered ontology, layering evidence cards, an authoritative metric ledger, and a claim graph into an always-current source of truth, not per-query RAG over raw chunks. A portable multi-agent "writer" runtime then composes a contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as_of <= T (no look-ahead); red-team verdicts flow back into the librarian. We evaluate on a self-collected, public corpus of 6,130 sources yielding 555,926 evidence cards (SEC EDGAR filings across 295 issuers and 11 sectors, U.S. Bureau of Labor Statistics releases, and Wikipedia). From the one library we compose four point-in-time reports on distinct theses and run eight reproducible experiments, whose headline metrics come from a deterministic quality-control gate, itself validated by defect-injection meta-evaluation at recall 1.0 and precision 1.0. A shared metric ledger removes 6,845 cross-section contradictions to zero. Tier-first selection is correct on 22/22 gold cases where a popularity-first baseline scores only 9/22; trust tiering leaks zero media-sourced numbers, and no government statistic displaces a company's own filing. A red-team refutation propagates back and self-corrects a later run with zero manual edits. Replay exhibits zero look-ahead violations across seven cutoffs while the library grows from 235,373 to 555,312 cards. Difficulty-tiered model routing exceeds the all-Opus quality ceiling while running 3.7x faster than serial.
Problem

Research questions and friction points this paper is trying to address.

fact drift
self-contradiction
provenance loss
evidence grounding
point-in-time consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

trust-tiered ontology
point-in-time knowledge
evidence-grounded writing
multi-agent writer
drift-free research
💼 Related Jobs
No related jobs found.
Xing Zhang
Xing Zhang
Amazon Visual Search / Binghamton University
Facial expression recognitionAugmented Reality (AR)computer visioncomputer graphicsmachine learning
Y
Yanwei Cui
AWS Generative AI Innovation Center
G
Guanghui Wang
AWS Generative AI Innovation Center
P
Peiyang He
AWS Generative AI Innovation Center