HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory
本文提出HERO框架,通过构建可追踪的异质记忆图和迭代图遍历结合人物档案的方法,解决长期代理记忆中信息丢失和语义漂移问题。
本文提出HERO框架,通过构建可追踪的异质记忆图和迭代图遍历结合人物档案的方法,解决长期代理记忆中信息丢失和语义漂移问题。
This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.
本文提出HERO框架,通过构建可追踪的异质记忆图和迭代图遍历结合人物档案的方法,解决长期代理记忆中信息丢失和语义漂移问题。
This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.