Towards Comprehensive Basketball Understanding

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建多模态基准BasketballBench和提出特定技能组合模型BasketballSkills,解决了篮球比赛中多种能力综合理解的问题。
📝 Abstract
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.
Problem

Research questions and friction points this paper is trying to address.

basketball understanding
event recognition
action localization
player identification
structured game knowledge
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal benchmark
BasketballSkills
domain-specific capabilities
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yirong Hu
School of Mathematical Sciences, Peking University, China
J
Jiayuan Rao
School of Artificial Intelligence, Shanghai Jiao Tong University, China
Y
Yu Zhang
School of Artificial Intelligence, Shanghai Jiao Tong University, China
Shangzhe Di
Shangzhe Di
Shanghai Jiao Tong University
Video UnderstandingMultimodal LearningComputer Vision
Weidi Xie
Weidi Xie
Shanghai Jiao Tong University | VGG, University of Oxford
Computer VisionAI for HealthcareAI for Science