Benchmarking Affordance Generalization with BusyBox
This work investigates the affordance generalization capabilities of vision–language–action (VLA) models when encountering novel objects that possess familiar physical properties but have never been seen before. To this end, we introduce BusyBox—a physical evaluation benchmark built upon six interchangeable modules, which enables the systematic and semi-automated assessment of model performance by generating visually diverse yet affordance-consistent object variants through module rotation and substitution. BusyBox provides, for the first time, a reproducible, low-cost, and easily constructible real-world testbed, accompanied by open-sourced 3D printing schematics, an electronic bill of materials, and a dual-arm robot demonstration dataset. Experiments reveal that current state-of-the-art open-source VLA models, such as π₀.₅ and GR00T-N1.6, exhibit limited generalization on this benchmark, thereby validating its effectiveness and necessity.