ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过ReactHuman基准测试解决多模态大语言模型在面对突发物理危险时的即时反应问题,使用高精度模拟和五项评估指标来衡量模型的合理、安全及物理基础性。
📝 Abstract
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing sudden household hazards; it spans 17 event families and over 1,000 bit-for-bit reproducible scenes with exact, annotation-free ground truth derived from 240 Hz rigid-body simulation, including adversarial objects whose appearance contradicts their physics (a foam anvil, a steel apple). We further design a five-metric suite that scores each reaction along three axes: reasonable, safe, and physically grounded. We physically execute every committed plan so that decisions have observable consequences. With this harness we evaluate seven representative MLLMs. Results show that reactive safety is far from solved: models mishandle roughly one hazard in three, act from fixed dispositions rather than the observed scene, trust appearance over motion, and miss interception points at meter scale even when the chosen action is correct; none of these failures shrink with model scale. ReactHuman thus offers both a fine-grained diagnosis and a scalable training signal toward physically grounded, safety-aware embodied agents. The benchmark can be found here: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled
Problem

Research questions and friction points this paper is trying to address.

reactive decision-making
embodied intelligence
safety-critical action
intuitive physics
multimodal large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

reactive decision-making
embodied intelligence
multimodal large language models
physical simulation
safety-critical actions
🔎 Similar Papers
Y
Yizhan Li
Université de Montréal, Mila – Quebec AI Institute
J
Jianxin You
Université de Montréal, Mila – Quebec AI Institute
M
Mengyang Xiong
McGill University, Mila – Quebec AI Institute
Y
Yinhuan Chen
McGill University
Zicheng Zhao
Zicheng Zhao
Nanjing University of Science and Technology
Knowledge GraphLarge Language ModelFew-shot LearningSemi-Supervised Learning
D
Dekun Wu
Université de Montréal, Mila – Quebec AI Institute
Dongqing Zhang
Dongqing Zhang
Chengdu University of Technology
transportation
Bang Liu
Bang Liu
Associate Professor at the University of Montreal, Canada CIFAR AI Chair at Mila
Natural Language ProcessingDeep LearningMachine LearningData Mining