AVA-Encoder: Towards Agent-Native Video Representation Learning
This work addresses the challenge that existing creative agents struggle to effectively learn from high-quality films, primarily due to the absence of structured video representations that are both content-faithful and amenable to agent-based reasoning and manipulation. To bridge this gap, the authors propose the AVA-Encoder framework, which introduces the first native video knowledge graph representation tailored for agent operation. This framework enables bidirectional conversion between video and knowledge graph through a self-encoding mechanism and incorporates a text-gradient optimization strategy guided by natural language update directions. The contributions include the first cinematic-scale knowledge graph dataset, a reconstruction benchmark, and a technical pipeline integrating multimodal asset linking, typed edge relations, and data-agnostic encoding. Experiments demonstrate that the method outperforms the strongest baseline by 20.7 percentage points in video reconstruction and achieves superior performance over hand-tuned strategies using 74.3% fewer prompts under a policy-only setting.