Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study
本文通过定义和分解多元同意指数Gamma,量化分析了多数投票在提高LLM答案准确性时可能失效的问题,并以GPT-4.1为例进行了研究。
本文通过定义和分解多元同意指数Gamma,量化分析了多数投票在提高LLM答案准确性时可能失效的问题,并以GPT-4.1为例进行了研究。
This work addresses the challenge of efficiently allocating computational resources during inference under a fixed budget to optimize test-time performance. It formalizes test-time scaling as a dynamic computation allocation problem and introduces a statistically guided routing strategy: within a two-stage verification framework, a lightweight verifier first processes an initial candidate set, and samples exhibiting high uncertainty or potential value are selectively routed to a stronger verification module. The approach integrates parameter-weighted token accounting with bootstrap-based significance testing. Evaluated across multiple mathematical and symbolic reasoning benchmarks, the method achieves a macro accuracy of 85.13% while reducing computational cost by 49.1%–58.9% compared to strong baselines, with negligible performance degradation.
This work addresses memory contamination in language agents caused by the retention of outdated information in long-term memory, which degrades decision accuracy. The authors propose TEPA, a novel mechanism that introduces revocable memory lifecycle management: observations are represented as keyed precedents, and conflicts between new evidence and existing memories are dynamically detected. Upon detecting such conflicts, TEPA revokes invalidated memories, ensuring retrieval is always grounded in the most current and valid knowledge. This approach enables dynamic falsification, auditability, and reactivation of memories. Evaluated across diverse memory drift scenarios, TEPA achieves an accuracy of 0.950, substantially outperforming conventional strategies that rely solely on append-only or overwrite-based memory updates.
This work addresses the degradation in retrieval performance and lack of decision credibility observed when frozen multimodal encoders are deployed compositionally due to varying connection paths. To tackle this, the authors propose CertBind, a novel framework that extends multimodal compositionality from representation learning to certifiable task-level decisions. CertBind introduces a four-tier certification mechanism—spanning nodes, edges, paths, and queries—integrated with anchored boundary modeling, contract-aware conformal ranking, overlap-aware budget allocation, and clean calibration to construct a certified retrieval system with a finite-sample recovery radius. Experiments demonstrate that under shared-path C-MCR settings, CertBind recovers 96.3% of the original retrieval performance while achieving perfect branch accuracy (1.000), effectively balancing compositional extensibility with decision reliability.
This study addresses the challenge of continuously reconstructing long-range vehicle trajectories from fixed highway cameras, which is hindered by perspective compression and scale attenuation. To this end, the authors introduce LoRFT, the first open benchmark specifically designed for this task, and propose Map-RSTNet, a map-aware sequence-to-sequence model. Map-RSTNet dynamically integrates local road structure within a road-geometry-aligned state space, leveraging a residual architecture and a geometric refresh mechanism to sustain trajectory continuity. Experimental results demonstrate that Map-RSTNet significantly outperforms existing methods on LoRFT, reducing Average Displacement Error (ADE), Final Displacement Error (FDE), and 5-second RMSE by 11.0%, 15.4%, and 10.5%, respectively, thereby effectively extending the usable length of trajectories captured by fixed cameras.
本文通过定义和分解多元同意指数Gamma,量化分析了多数投票在提高LLM答案准确性时可能失效的问题,并以GPT-4.1为例进行了研究。
This work addresses the challenge of efficiently allocating computational resources during inference under a fixed budget to optimize test-time performance. It formalizes test-time scaling as a dynamic computation allocation problem and introduces a statistically guided routing strategy: within a two-stage verification framework, a lightweight verifier first processes an initial candidate set, and samples exhibiting high uncertainty or potential value are selectively routed to a stronger verification module. The approach integrates parameter-weighted token accounting with bootstrap-based significance testing. Evaluated across multiple mathematical and symbolic reasoning benchmarks, the method achieves a macro accuracy of 85.13% while reducing computational cost by 49.1%–58.9% compared to strong baselines, with negligible performance degradation.
This work addresses memory contamination in language agents caused by the retention of outdated information in long-term memory, which degrades decision accuracy. The authors propose TEPA, a novel mechanism that introduces revocable memory lifecycle management: observations are represented as keyed precedents, and conflicts between new evidence and existing memories are dynamically detected. Upon detecting such conflicts, TEPA revokes invalidated memories, ensuring retrieval is always grounded in the most current and valid knowledge. This approach enables dynamic falsification, auditability, and reactivation of memories. Evaluated across diverse memory drift scenarios, TEPA achieves an accuracy of 0.950, substantially outperforming conventional strategies that rely solely on append-only or overwrite-based memory updates.
This work addresses the degradation in retrieval performance and lack of decision credibility observed when frozen multimodal encoders are deployed compositionally due to varying connection paths. To tackle this, the authors propose CertBind, a novel framework that extends multimodal compositionality from representation learning to certifiable task-level decisions. CertBind introduces a four-tier certification mechanism—spanning nodes, edges, paths, and queries—integrated with anchored boundary modeling, contract-aware conformal ranking, overlap-aware budget allocation, and clean calibration to construct a certified retrieval system with a finite-sample recovery radius. Experiments demonstrate that under shared-path C-MCR settings, CertBind recovers 96.3% of the original retrieval performance while achieving perfect branch accuracy (1.000), effectively balancing compositional extensibility with decision reliability.
This study addresses the challenge of continuously reconstructing long-range vehicle trajectories from fixed highway cameras, which is hindered by perspective compression and scale attenuation. To this end, the authors introduce LoRFT, the first open benchmark specifically designed for this task, and propose Map-RSTNet, a map-aware sequence-to-sequence model. Map-RSTNet dynamically integrates local road structure within a road-geometry-aligned state space, leveraging a residual architecture and a geometric refresh mechanism to sustain trajectory continuity. Experimental results demonstrate that Map-RSTNet significantly outperforms existing methods on LoRFT, reducing Average Displacement Error (ADE), Final Displacement Error (FDE), and 5-second RMSE by 11.0%, 15.4%, and 10.5%, respectively, thereby effectively extending the usable length of trajectories captured by fixed cameras.