Institution profile

ANTLAB

Industry research
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs

May 06, 2026

This work addresses the challenge of automatically generating trustworthy multi-hop reasoning questions from scientific literature, where relationships among multimodal elements are often implicit and difficult to verify. To this end, it introduces knowledge graphs into a self-play framework for scientific documents, constructing a unified graph to generate multi-hop relational questions and providing verifiable reward signals grounded in structured factual knowledge. By leveraging an information asymmetry mechanism, a single small-scale vision-language model alternately assumes the roles of questioner and answerer during training. The proposed approach significantly outperforms text-only self-play baselines on both public benchmarks and a newly curated cross-document multi-hop question answering dataset, with performance gains becoming more pronounced as the number of reasoning hops increases.

0 citationsRead paper
Recent publications

Latest Papers

SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs

May 06, 2026

This work addresses the challenge of automatically generating trustworthy multi-hop reasoning questions from scientific literature, where relationships among multimodal elements are often implicit and difficult to verify. To this end, it introduces knowledge graphs into a self-play framework for scientific documents, constructing a unified graph to generate multi-hop relational questions and providing verifiable reward signals grounded in structured factual knowledge. By leveraging an information asymmetry mechanism, a single small-scale vision-language model alternately assumes the roles of questioner and answerer during training. The proposed approach significantly outperforms text-only self-play baselines on both public benchmarks and a newly curated cross-document multi-hop question answering dataset, with performance gains becoming more pronounced as the number of reasoning hops increases.

0 citationsRead paper