Score
Constructs perceptual models (e.g., psychoacoustic models) to relate physical stimuli to human perception, producing models and metrics for perceptual evaluation and system tuning.
This study investigates whether speech language models (SLMs) possess the human-like capacity for sound symbolism—the perceptual association between speech sounds and sensory attributes such as roundedness or sharpness. Departing from prior text- or image-based approaches, this work presents the first systematic evaluation of SLMs’ sound symbolic judgments using authentic human speech recordings across auditory, cross-modal, and visual dimensions, benchmarked against human behavioral data. Through acoustic feature analysis—particularly spectral tilt—and controlled visual experiments, the study reveals that SLMs significantly diverge from human perception in auditory judgments, fail to capture the key acoustic cues underlying sound symbolism, and cannot reliably map sounds to corresponding shapes, thereby exposing critical limitations in their speech representations.
This study investigates whether vision-language models (VLMs) can effectively emulate human judgments in perceptual image quality assessment, potentially replacing costly psychophysical experiments. For the first time, we systematically evaluate six VLMs—four closed-source and two open-source—against human judgments across three dimensions: contrast, color saturation, and overall preference, using established psychophysical data as a benchmark. Our analysis integrates attribute-weighted evaluation and intra-model consistency metrics. Results show that VLMs achieve up to 0.93 correlation with human judgments on color saturation but perform notably weaker on contrast. Most models align with human behavior by prioritizing color saturation in overall preference. A key contribution is revealing the trade-off between model self-consistency and human alignment, and demonstrating that enhancing perceptual separability improves human-model agreement.
This study addresses the limited capability of large audio language models in fundamental auditory perception—such as pitch, loudness, and spatial location—where their performance often approaches random guessing and fails to match human superiority in comparative tasks. To systematically evaluate this gap, the authors propose SonicBench, the first psychophysics-based benchmark integrating controllable audio generation, a dual-task paradigm of identification and comparison, linear probing analysis, and controlled experiments. Findings reveal that frozen audio encoders already capture physical auditory cues effectively (accuracy ≥60%), indicating that the primary bottleneck lies not in perceptual encoding but in subsequent alignment and decoding stages. This work provides the first systematic characterization of the boundaries of audio language models in basic auditory perception.
To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.
This study investigates the human-like capability of large language models (LLMs) in multimodal perceptual intensity modeling, benchmarking against human cross-sensory (e.g., visual–auditory–tactile) intensity ratings as ground truth. Method: We introduce the first perceptual-intensity-based quantitative benchmark and a contrastive evaluation framework integrating quantitative correlation analysis with qualitative error-pattern mining. Contribution/Results: GPT-4 and GPT-4o significantly outperform GPT-3.5 and GPT-4o-mini, yet GPT-4o does not surpass GPT-4—suggesting current multimodal fusion fails to enhance embodied perceptual grounding. Models consistently exhibit non-human reasoning patterns, including multisensory overestimation and reliance on superficial semantic associations. This work pioneers the integration of perceptual intensity into LLM multimodal evaluation, establishing a novel paradigm and reproducible benchmark for semantic grounding and embodied intelligence research.
This work addresses the limited intuitive understanding of physical dynamics in current vision-language models, which hinders their ability to generalize physical rules to novel scenarios. It presents the first systematic exploration of leveraging reinforcement learning–driven interactive training to enable pretrained vision-language models to learn physical dynamics within simulated environments, with the goal of acquiring generalizable physical intuition. This approach challenges the prevailing paradigm of static fine-tuning. Experimental results demonstrate that while the models exhibit improved performance on in-distribution tasks, their generalization to new tasks with similar visual and physical characteristics remains limited. These findings suggest that interactive learning alone is insufficient for developing robust, transferable physical intuition, thereby highlighting critical directions for future research.
Human tactile perception relies on multimodal cues, yet the mapping between low-level tactile signals and high-level perceptual representations remains poorly understood, hindering the application of tactile technologies in digital and robotic systems. This work proposes an interpretable computational framework that, for the first time, systematically integrates thermal conduction and deformation (compliance) cues across pressing, static contact, and sliding interactions. The framework comprises three interconnected models that map physical interaction features to psychophysical perceptual attributes and enable material classification. By combining multimodal tactile sensing, psychophysical modeling, and interpretable machine learning, the approach significantly improves material identification accuracy and reveals the critical roles of thermal and compliance cues in perceptual modeling and multimodal cue integration.
This work addresses the lack of intuitive, immediate feedback on fitting errors in existing model-fitting approaches. It proposes an interactive fitting framework that integrates visual and auditory feedback: as users manipulate parametric curves, the system synthesizes audio in real time, with greater model-data discrepancies producing louder and more dissonant sounds. This is the first approach to incorporate auditory cues into model exploration, enabling multisensory assessment of fit quality. Combining interactive visualization, real-time audio synthesis, and Gaussian process regression, the method demonstrates effectiveness and generalizability across four diverse case studies—golf putting, dilution experiments, cosmological parameter estimation, and temperature data fitting—significantly enhancing users’ intuitive perception of model misfit.
This study addresses a central challenge in industrial design: effectively translating consumers’ abstract and subjective aesthetic preferences—such as “sportiness”—into actionable design guidance. To this end, the authors propose a human-centered computing framework that, for the first time, integrates subjective evaluations from both consumers and designers, domain-specific design features (e.g., wheel spoke configurations), and objective visual metrics extracted via computer vision (e.g., texture). By jointly modeling these multi-source features, the approach constructs an interpretable aesthetic perception model that explicitly links aesthetic judgments to concrete design elements. This enables product teams to accurately anticipate user preferences early in the design process, facilitating efficient exploration and critical iteration of design alternatives.
Existing large models struggle to generate non-obvious yet physically feasible tool-use strategies in open-world settings due to insufficient grounding in visual and physical constraints. To address this, this work introduces MM-CreativityBench, the first benchmark for systematically evaluating embodied creativity in large models. The proposed approach incorporates a functional attribute alignment mechanism that leverages a functional knowledge base for supervision, multi-view structured scene representations, multi-turn interactive reasoning, and preference learning via Direct Preference Optimization. This framework significantly improves the model’s accuracy in selecting appropriate objects and parts, effectively mitigates hallucination and grounding errors, and consistently enhances performance across creative physical reasoning tasks.