QuantWAMs: Calibrating at the Right Granularity for World Action Models

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently deploying world action models (WAMs) with existing post-training quantization methods, which often rely on open-loop objectives, homogeneous assumptions, and calibration distributions misaligned with real-world deployment conditions. To overcome these limitations, we propose QuantWAMs—the first post-training quantization framework tailored for WAMs—leveraging a fine-grained calibration strategy that integrates model architecture, trajectory distribution, and task-specific objectives. Key innovations include shared-base outlier calibration, joint objective-aware saliency weighting, fixed-intervention trajectory auditing, empirical Fisher-based saliency analysis, coordinate-compatible activation pooling, and closed-loop reachable-state-guided denoising scheduling. Under W4A4 quantization, our method incurs only a 0.2–0.7 percentage point performance drop in simulated tasks and successfully executes three real-world robotic manipulation tasks, achieving a peak memory footprint of 29% relative to FP16 and module-level speedups of 1.4–1.6×.
📝 Abstract
World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-training quantization (PTQ) methods are poorly suited to WAMs because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment. We present QuantWAMs, a PTQ framework that aligns quantization decisions with the calibration context defined by model structure, rollout distribution, and task objective. QuantWAMs introduces three strategies: shared-basis outlier calibration, which pools activation evidence only across coordinate-compatible modules; co-training-objective saliency, which computes empirical-Fisher scores from the joint video--action gradient and assigns weight precision at a calibration-stable layer granularity; and fixed-intervention rollout auditing, which revises denoising-step protection schedules using reachable closed-loop states without changing the precision budget. We evaluate QuantWAMs on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and real-robot manipulation with an AgiBot G2. Under a W4A4-dominant setting, the reported simulation means differ from FP16 by 0.2--0.7 percentage points. Real-robot trials further establish deployment feasibility on three manipulation tasks. For the targeted video and action blocks, QuantWAMs reduces peak weight-and-activation memory to about 29\% of FP16 and provides 1.4--1.6$\times$ block-level speedups.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
post-training quantization
closed-loop execution
calibration distribution
efficient deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

QuantWAMs
post-training quantization
World Action Models
closed-loop deployment
calibration granularity
🔎 Similar Papers
No similar papers found.
J
Jiacheng Zhou
College of Intelligent Robotics and Advanced Manufacturing, Fudan University
J
Jinfan Lv
College of Intelligent Robotics and Advanced Manufacturing, Fudan University
R
Ruixuan Li
College of Intelligent Robotics and Advanced Manufacturing, Fudan University
L
Longtai Zhang
College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Yan Wang
Yan Wang
Professor in East China Normal University
computer visionmedical image analysis
Wenqiang Zhang
Wenqiang Zhang
School of Computer Science, Fudan University
RoboticMedical ImageComputer vision
L
Lizhe Qi
College of Intelligent Robotics and Advanced Manufacturing, Fudan University