Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
研究解决了全双工语音LLM中不适当发言的问题,通过因果分析确定原因,并提出一种实时干预方法有效抑制了这种现象,同时保持了正常响应。
研究解决了全双工语音LLM中不适当发言的问题,通过因果分析确定原因,并提出一种实时干预方法有效抑制了这种现象,同时保持了正常响应。
This work identifies and formalizes the phenomenon of “physical misgeneralization,” wherein generative sequential models in physical environments produce global distributions of conserved quantities—such as path length or mechanical energy—that deviate from design intent due to accumulated local modeling errors. The authors introduce controlled synthetic tasks to elucidate the underlying mechanism: local inaccuracies propagate through physical constraints, distorting the global distribution of these quantities. To address this, they propose a data bias kernel that predicts the direction of distributional shift, enabling a structured intervention strategy informed by this prediction. Experiments on maze navigation and double-pendulum dynamics demonstrate that the method effectively anticipates and mitigates distortions in physical quantity distributions, thereby validating both the proposed mechanistic explanation and the efficacy of the intervention approach.
This work addresses the challenge of optimizing three-layer routing in telecommunications networks to achieve low latency and high resilience under dual-link failures. The problem is formulated as a graph optimization task of finding vertex-disjoint paths. For the first time, the Quantum Approximate Optimization Algorithm (QAOA) is innovatively applied to this domain, with a unified objective function designed to jointly minimize latency and ensure fault tolerance against dual-link failures. Feasibility is demonstrated through experiments on both quantum simulators and real hardware. Evaluations on a 5-node, 7-edge topology—covering both independent and highly correlated failure scenarios—consistently yield optimal solutions, evidenced by the lowest energy states and highest sampling frequencies, thereby validating the effectiveness and practicality of the proposed approach.
This work addresses the challenge of multi-hop question answering over hybrid table-text data, where retrieval noise significantly degrades performance and conventional RAG approaches struggle to model reasoning chains due to their flat document representation. The authors propose the first zero-shot graph-based framework that leverages large language models to dynamically construct an evidence graph from noisy retrieval results—treating documents as nodes and semantic relations as edges—to explicitly capture multi-hop reasoning paths and automatically identify bridging documents. Requiring no task-specific training, the method achieves 48.80 EM on OTT-QA, outperforming strong baselines by 19.9 points, matching the performance of fine-tuned retrieval models (CORE: 49.0 EM), and approaching that of the state-of-the-art system (COS: 56.9 EM).
Existing recommender systems often oversimplify timestamps as numerical values or periodic signals, neglecting the dynamic influence of geotemporal context—such as holidays, local events, and seasonal patterns—on user behavior. To address this, we propose a lightweight, LLM-driven spatiotemporal contextual embedding method: leveraging large language models, it jointly processes timestamps and coarse-grained geographic information to generate semantically rich embeddings encoding holiday observances, seasonal trends, and locale-specific events. We design an embedding informativeness metric and integrate it into a sequential recommendation framework, enabling adaptive feature fusion and auxiliary loss optimization. Extensive experiments on multiple benchmark datasets demonstrate significant improvements in recommendation accuracy. Furthermore, we release a high-quality, spatiotemporally enhanced version of the MovieLens dataset to advance research in context-aware recommendation.
研究解决了全双工语音LLM中不适当发言的问题,通过因果分析确定原因,并提出一种实时干预方法有效抑制了这种现象,同时保持了正常响应。
This work identifies and formalizes the phenomenon of “physical misgeneralization,” wherein generative sequential models in physical environments produce global distributions of conserved quantities—such as path length or mechanical energy—that deviate from design intent due to accumulated local modeling errors. The authors introduce controlled synthetic tasks to elucidate the underlying mechanism: local inaccuracies propagate through physical constraints, distorting the global distribution of these quantities. To address this, they propose a data bias kernel that predicts the direction of distributional shift, enabling a structured intervention strategy informed by this prediction. Experiments on maze navigation and double-pendulum dynamics demonstrate that the method effectively anticipates and mitigates distortions in physical quantity distributions, thereby validating both the proposed mechanistic explanation and the efficacy of the intervention approach.
This work addresses the challenge of optimizing three-layer routing in telecommunications networks to achieve low latency and high resilience under dual-link failures. The problem is formulated as a graph optimization task of finding vertex-disjoint paths. For the first time, the Quantum Approximate Optimization Algorithm (QAOA) is innovatively applied to this domain, with a unified objective function designed to jointly minimize latency and ensure fault tolerance against dual-link failures. Feasibility is demonstrated through experiments on both quantum simulators and real hardware. Evaluations on a 5-node, 7-edge topology—covering both independent and highly correlated failure scenarios—consistently yield optimal solutions, evidenced by the lowest energy states and highest sampling frequencies, thereby validating the effectiveness and practicality of the proposed approach.
This work addresses the challenge of multi-hop question answering over hybrid table-text data, where retrieval noise significantly degrades performance and conventional RAG approaches struggle to model reasoning chains due to their flat document representation. The authors propose the first zero-shot graph-based framework that leverages large language models to dynamically construct an evidence graph from noisy retrieval results—treating documents as nodes and semantic relations as edges—to explicitly capture multi-hop reasoning paths and automatically identify bridging documents. Requiring no task-specific training, the method achieves 48.80 EM on OTT-QA, outperforming strong baselines by 19.9 points, matching the performance of fine-tuned retrieval models (CORE: 49.0 EM), and approaching that of the state-of-the-art system (COS: 56.9 EM).
Existing recommender systems often oversimplify timestamps as numerical values or periodic signals, neglecting the dynamic influence of geotemporal context—such as holidays, local events, and seasonal patterns—on user behavior. To address this, we propose a lightweight, LLM-driven spatiotemporal contextual embedding method: leveraging large language models, it jointly processes timestamps and coarse-grained geographic information to generate semantically rich embeddings encoding holiday observances, seasonal trends, and locale-specific events. We design an embedding informativeness metric and integrate it into a sequential recommendation framework, enabling adaptive feature fusion and auxiliary loss optimization. Extensive experiments on multiple benchmark datasets demonstrate significant improvements in recommendation accuracy. Furthermore, we release a high-quality, spatiotemporally enhanced version of the MovieLens dataset to advance research in context-aware recommendation.