From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish
研究探讨频率和 surprisal 在西班牙语词汇习得和阅读中的作用,发现频率对儿童词汇习得更重要,而 surprisal 更能预测成人阅读时的处理难度。
研究探讨频率和 surprisal 在西班牙语词汇习得和阅读中的作用,发现频率对儿童词汇习得更重要,而 surprisal 更能预测成人阅读时的处理难度。
为解决第三方服务商提供的轨迹数据查询结果不可信问题,提出VTRQ框架,通过空间和时间认证数据结构提高查询与验证效率。
This study evaluates the biological fidelity of the self-supervised audiovisual speech model AV-HuBERT in modeling human multisensory integration, particularly the McGurk effect. Using a standard McGurk experimental paradigm and statistical analyses, the authors compare AV-HuBERT’s responses to incongruent audiovisual stimuli against behavioral data from 44 human participants. The model exhibits a striking alignment with human perception in auditory dominance rates (32.0% vs. 31.8%) but demonstrates significantly higher phoneme fusion rates (68.0% vs. 47.7%). Moreover, it lacks the variability and diversity of perceptual errors characteristic of human responses. These findings indicate that while AV-HuBERT captures certain aspects of human multisensory speech integration, it remains limited by rigid decision-making mechanisms that fail to replicate the flexibility and stochasticity inherent in human perception.
This study investigates the capacity of large language models (LLMs) to detect, reproduce, and correct authentic speech errors produced by native Spanish speakers, thereby exposing fundamental limitations in their emulation of human linguistic cognition. Methodologically, it integrates interdisciplinary approaches: (1) constructing a curated, annotated corpus of 500+ naturally occurring Spanish errors; (2) conducting high-temporal-resolution EEG experiments to characterize real-time neural processing dynamics; and (3) systematically evaluating state-of-the-art LLMs—including GPT and Gemini—on error detection, generative generalization, and fault-tolerant reasoning. This work achieves the first tri-level, cross-domain modeling of linguistic errors across cognitive representation, neurophysiological response, and AI behavioral output. Results advance theoretical understanding of Spanish linguistic competence and variation, while providing an empirical foundation and methodological framework for developing cognitively aligned, robust, and error-resilient NLP systems.
This paper addresses the core challenge in multi-to-multi bond trading platforms where dealers cannot observe counterparties’ quotes and struggle to ensure profitability. We propose the first general analytical framework integrating causal inference with probabilistic graphical models. Methodologically, we distinguish generative versus discriminative modeling approaches, explicitly capture the dynamic impact of RFQ (Request-for-Quote) negotiation mechanisms, identify key pricing drivers via causal interventions, and design prediction evaluation metrics tailored for optimal pricing. Our contribution lies in the first incorporation of structured causal modeling into electronic RFQ decision-making—overcoming the limitations of traditional black-box predictive models. Empirical results demonstrate significant improvements in price prediction accuracy and revenue estimation reliability, thereby enhancing dealers’ pricing capability and profitability under information asymmetry.
研究探讨频率和 surprisal 在西班牙语词汇习得和阅读中的作用,发现频率对儿童词汇习得更重要,而 surprisal 更能预测成人阅读时的处理难度。
为解决第三方服务商提供的轨迹数据查询结果不可信问题,提出VTRQ框架,通过空间和时间认证数据结构提高查询与验证效率。
This study evaluates the biological fidelity of the self-supervised audiovisual speech model AV-HuBERT in modeling human multisensory integration, particularly the McGurk effect. Using a standard McGurk experimental paradigm and statistical analyses, the authors compare AV-HuBERT’s responses to incongruent audiovisual stimuli against behavioral data from 44 human participants. The model exhibits a striking alignment with human perception in auditory dominance rates (32.0% vs. 31.8%) but demonstrates significantly higher phoneme fusion rates (68.0% vs. 47.7%). Moreover, it lacks the variability and diversity of perceptual errors characteristic of human responses. These findings indicate that while AV-HuBERT captures certain aspects of human multisensory speech integration, it remains limited by rigid decision-making mechanisms that fail to replicate the flexibility and stochasticity inherent in human perception.
This study investigates the capacity of large language models (LLMs) to detect, reproduce, and correct authentic speech errors produced by native Spanish speakers, thereby exposing fundamental limitations in their emulation of human linguistic cognition. Methodologically, it integrates interdisciplinary approaches: (1) constructing a curated, annotated corpus of 500+ naturally occurring Spanish errors; (2) conducting high-temporal-resolution EEG experiments to characterize real-time neural processing dynamics; and (3) systematically evaluating state-of-the-art LLMs—including GPT and Gemini—on error detection, generative generalization, and fault-tolerant reasoning. This work achieves the first tri-level, cross-domain modeling of linguistic errors across cognitive representation, neurophysiological response, and AI behavioral output. Results advance theoretical understanding of Spanish linguistic competence and variation, while providing an empirical foundation and methodological framework for developing cognitively aligned, robust, and error-resilient NLP systems.
This paper addresses the core challenge in multi-to-multi bond trading platforms where dealers cannot observe counterparties’ quotes and struggle to ensure profitability. We propose the first general analytical framework integrating causal inference with probabilistic graphical models. Methodologically, we distinguish generative versus discriminative modeling approaches, explicitly capture the dynamic impact of RFQ (Request-for-Quote) negotiation mechanisms, identify key pricing drivers via causal interventions, and design prediction evaluation metrics tailored for optimal pricing. Our contribution lies in the first incorporation of structured causal modeling into electronic RFQ decision-making—overcoming the limitations of traditional black-box predictive models. Empirical results demonstrate significant improvements in price prediction accuracy and revenue estimation reliability, thereby enhancing dealers’ pricing capability and profitability under information asymmetry.