DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge posed by the disparity between simple and complex questions in medical video question answering by proposing a difficulty-aware multimodal evidence planning approach. The method integrates caption, visual, and procedural context evidence to rank candidate videos, then generates and refines temporal segments through high-scoring anchor expansion. Its key innovation lies in transforming explicit question difficulty labels into dynamic control strategies during inference, enabling adaptive adjustment of modality weights, evidence aggregation schemes, boundary thresholds, and reranking intensity. Evaluated on the NLPCC 2026 Shared Task, the proposed model achieved first place with an average score of 0.2728. Ablation studies further demonstrate that the introduced mechanisms yield particularly significant performance gains on complex questions.
📝 Abstract
We describe DAEP, team BIGC's submission to NLPCC 2026 Shared Task 1 Track 3: Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The task requires retrieving the target video from 50 candidates and localizing the answer-supporting span. DAEP ranks videos with subtitle, visual, and procedural-context evidence, expands high-scoring anchors into temporal spans, and reranks spans for final output. Its main design is to convert the task-provided simple/complex input label into an inference-time evidence plan controlling modality weights, Top-K aggregation, boundary threshold, expansion length, and reranking strength. In the official evaluation, BIGC ranks first among ten systems with an Average score of 0.2728. Validation ablations show that visual evidence, procedural context, and difficulty-aware planning improve ranking quality, with the largest gain on complex questions.
Problem

Research questions and friction points this paper is trying to address.

Temporal Answer Grounding
Video Corpus
Difficulty-Aware
Medical Video
Answer Span Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Difficulty-Aware Planning
Temporal Answer Grounding
Multimodal Evidence Fusion
Video Corpus Retrieval
Adaptive Inference Strategy
T
Tianjian He
TikTok, ByteDance, China
Y
Yujie Liu
Beijing Institute of Graphic Communication, Beijing, China
Z
Zhiping Huang
Lingnan University, Hong Kong, China
C
Changbo Xu
Beijing Institute of Graphic Communication, Beijing, China