TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in existing tennis video understanding methods, which struggle to bridge fine-grained stroke perception with high-level tactical reasoning due to a lack of stroke-level evidence in decision modeling. To overcome this, we propose a stroke-evidence-driven multimodal large language model that parses rallies into structured sequences of stroke events and incorporates a tactics-graph-guided temporal reasoning mechanism, enabling end-to-end inference across events, relationships, evidence, and tactics. We introduce TRACE, the first large-scale expert-annotated benchmark, featuring stroke attributes, inter-stroke relations, hierarchical tactical labels, and evidence-anchored question answering. Evaluated on 11,189 rallies from TRACE, our model significantly improves accuracy and interpretability across open-ended answer prediction, tactic classification, stroke sequence extraction, and key action localization tasks.
📝 Abstract
Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling their tactical dependencies or generate high-level analyses without grounding them in the underlying events. To bridge this perception-to-understanding gap, we formulate stroke-evidence-grounded tactical reasoning, a new rally-level task that requires models to jointly predict an open-ended answer, a hierarchical tactic label, an ordered sequence of supporting strokes, and decisive key actions, with each evidence stroke anchored to its racket-ball contact frame. We further introduce TRACE (Tactical Reasoning with Action-Chain Evidence in Tennis), a large-scale expert-annotated benchmark containing 11,189 rally videos, 41,485 stroke events, 25,429 tactical units, and 11,189 question-answer pairs, which unifies fine-grained stroke attributes, cross-stroke tactical relations, hierarchical tactic annotations, and evidence-grounded questions across factual perception, tactical understanding, and decision reasoning. Building on TRACE, we propose TennisVAR (Tennis Video Action-chain Reasoner), an evidence-grounded multimodal large language model that follows an "event-relation-evidence-tactic" reasoning paradigm, where an Event Parsing Module converts continuous rallies into explicit stroke-event sequences while a Tactical Graph-Guided Temporal Reasoner jointly models rally progression and same-player decision dependencies to identify question-relevant evidence and decisive actions.
Problem

Research questions and friction points this paper is trying to address.

tactical reasoning
stroke evidence
tennis video understanding
multimodal reasoning
action-chain grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

stroke-evidence-grounded reasoning
multimodal large language model
tactical reasoning
action-chain evidence
event parsing
Y
Yifan Mei
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
Q
Qingling Shi
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University
C
Changli Wu
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University; Shanghai Innovation Institute
J
Jiayuan Rao
School of Artificial Intelligence, Shanghai Jiao Tong University
Jiayi Ji
Jiayi Ji
Rutgers University
L
Liujuan Cao
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University