EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决UAV视频问答中相机运动与场景变化分离问题,提出EgoSIS方法,通过三阶段将RGB导出的双向流转换为规范化的视觉证据。
📝 Abstract
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. Factorized Visual Ego-Transitions (FVET) fits a robust image-plane transition and exposes motion, residual-support, and reliability factors. Reliability-Gated Ego-Transition Memory (ReTEM) uses reliability-weighted updates for a bounded history and re-anchors it at cuts or sustained uncertainty. Ego-Aligned Spatial Evidence (EASE) warps supported visual features into each segment's local anchor and injects four spatial evidence tokens per visual slice through zero-initialized residuals, without changing Qwen's visual-token count. On SIS-Bench, EgoSIS-8B obtains 89.9\% perception, 82.5\% perception-plus-memory, and 76.2\% overall accuracy, with the largest gains concentrated in self-awareness perception and memory. The adapter thus provides an interpretable interface between optical flow and spatial reasoning.
Problem

Research questions and friction points this paper is trying to address.

UAV
video question answering
camera motion
scene changes
RGB-only multimodal models
Innovation

Methods, ideas, or system contributions that make the work stand out.

EgoSIS
Factorized Visual Ego-Transitions
Reliability-Gated Ego-Transition Memory
Ego-Aligned Spatial Evidence
🔎 Similar Papers
No similar papers found.
J
Jingpu Yang
Beihang University, Beijing, China
Fengxian Ji
Fengxian Ji
Northeast University
agent、Machine learnin、CV
M
Mingxuan Cui
Northeastern University, Shenyang, China
Y
Yilin Sun
Beihang University, Beijing, China
Hang Zhang
Hang Zhang
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences, Beijing 100094, China
J
Jianhua Zhu
Beihang University, Beijing, China
Y
Yufeng Wang
Beihang University, Beijing, China