Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

πŸ“… 2026-08-14
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of insufficient local evidence and score bias in multimodal large language models (MLLMs) for zero-shot micro-gesture recognition. We propose a test-time evidence calibration framework that introduces a novel calibration mechanism leveraging tree search to acquire fine-grained visual evidence. Combined with a multi-cue fusion strategy, this approach effectively mitigates motion perception bottlenecks and enhances prediction reliability. Experimental results demonstrate that the proposed framework achieves accuracies of 26.84% on iMiGUE and 22.10% on MA-52, significantly outperforming the Qwen2.5-VL baseline. These findings validate the method’s effectiveness in refining reasoning granularity and overcoming critical obstacles in MLLM-based micro-action recognition.
πŸ“ Abstract
While Multimodal Large Language Models (MLLMs) excel in general video understanding, their capability in fine-grained and motion-centric tasks remains limited. This limitation is particularly critical in micro-gesture recognition (MGR), where micro-gestures (MGs) - subtle, short-duration, and spatially localized human movements - serve as key discriminative signals for implicit affective analysis, yet are easily neglected following common prompting practices. Although MGR has been intensively studied by many discriminative approaches, the use of MLLMs for MGR is underexplored, with notably poor performance. We hypothesize that the motion-sensitive representation ability of MLLMs is constrained by their inherent single-pass forward inference, which can be substantially enhanced through carefully designed test-time guidance. Motivated by this, building on our prior findings regarding temporal insensitivity in Video LLMs, we diagnose zero-shot MGR errors in the Negative Log-Likelihood (NLL) space. We observe that MLLMs suffer from two bottlenecks: 1) insufficient localized evidence and 2) severe score biases driven by language and motion-agnostic appearances. Thus, we propose a novel test-time evidence calibration framework that improves both reasoning details and prediction reliability. Specifically, we introduce a tree search mechanism to progressively acquire localized, fine-grained visual evidence, coupled with a test-time calibration module to mitigate score biases. The multi-cue fusion module then integrates evidence from multiple cues without relying on a single cue for final prediction. Our framework achieves mean-class accuracies of 26.84\% on iMiGUE and 22.10\% on MA-52, significantly outperforming the Qwen2.5-VL baseline, which produces 16.15\% and 10.20\%, respectively. The code will be available at https://zero-melo.github.io/Zero-MELO.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Micro-Gesture Recognition
Multimodal Large Language Models
Test-Time Evidence Calibration
Score Biases
Localized Evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Evidence Calibration
Tree Search Mechanism
Zero-Shot Micro-Gesture Recognition
Multimodal LLMs
Score Bias Mitigation
πŸ’Ό Related Jobs
No related jobs found.