Certified decoding of quantum LDPC codes
该研究解决了量子LDPC码的解码难题,通过构建概率模型并开发两种新型解码器,一种基于采样提供最优性证明,另一种基于区域实现精确最大似然解码。
该研究解决了量子LDPC码的解码难题,通过构建概率模型并开发两种新型解码器,一种基于采样提供最优性证明,另一种基于区域实现精确最大似然解码。
This study addresses the bias in traditional ATT estimation arising from parallel trends violations in subsets of the treatment group by proposing a Credible Subgroup Local ATT. This method achieves robust causal identification under weak parallel trends assumptions through reweighting and honest sensitivity bounds, point-identifying a novel estimand based solely on cohort-specific parallel trends at the cost of a narrowed estimation scope. Simulations confirm the method’s validity, while an empirical application demonstrates that the previously documented positive effect of the shale gas boom on housing prices is actually driven by trend bias rather than genuine causality. By revealing no true causal impact through credible subgroup analysis, these findings underscore the approach’s critical value in correcting selection bias when standard identifying assumptions are compromised.
This work addresses the lack of a unified configuration governance mechanism in heterogeneous multi-agent systems, which hinders versioned, auditable, and cross-framework consistent management. The authors propose a framework-agnostic reference model for agent configuration governance that normalizes diverse configurations into a canonical configuration graph via semantic projection and enforces uniform governance semantics over this graph. Key innovations include typed and independently versioned configuration items, strict decoupling of configuration from runtime, a lattice-based monotonic influence propagation mechanism, and dependency-aware immutable revisions with provenance tracking. The model is validated across LangGraph, CrewAI, and OpenAI Agents SDK, demonstrating governance-equivalent ACM representations across 27 governance scenarios and 9 propagation cases, while guaranteeing convergence, termination, and a unique fixed point.
This study addresses the lack of systematic evaluation of Programmatic Tool Calling (PTC) versus conventional JSON-based tool invocation under realistic task conditions. Conducting an empirical comparison across 14 prominent language models on the BFCL v4 benchmark, this work provides the first comprehensive validation of PTC’s superiority and robustness across diverse models and scenarios. The methodology exposes tools as typed Python stubs, enabling models to generate executable scripts that support both chained and parallel invocations. Results demonstrate that 11 out of 14 models either maintain or significantly improve performance under PTC, with the GPT-5.6 series achieving a 10.6% gain. Notably, PTC sustains its advantage even in complex settings involving parallel tool use and contextual interference.
This work investigates the dual anchoring phenomenon in vision-language models, wherein model outputs are jointly influenced by textual and visual contexts—even when such contexts are irrelevant or incorrect yet remain plausible in the real world. To this end, we introduce ENTRAP-VL, a novel evaluation framework that establishes the first context-anchoring taxonomy tailored for multimodal settings, explicitly distinguishing between text-driven and vision-driven anchoring mechanisms and incorporating a “scene-inconsistent but world-plausible” realism dimension. Guided by a two-axis design (relevance × realism), we construct a structured probe dataset comprising 1,500 samples across eight scene categories, along with a standardized evaluation protocol. The ENTRAP-VL dataset and associated tools are publicly released to provide the community with a reproducible and extensible foundation for systematic research.
该研究解决了量子LDPC码的解码难题,通过构建概率模型并开发两种新型解码器,一种基于采样提供最优性证明,另一种基于区域实现精确最大似然解码。
This study addresses the bias in traditional ATT estimation arising from parallel trends violations in subsets of the treatment group by proposing a Credible Subgroup Local ATT. This method achieves robust causal identification under weak parallel trends assumptions through reweighting and honest sensitivity bounds, point-identifying a novel estimand based solely on cohort-specific parallel trends at the cost of a narrowed estimation scope. Simulations confirm the method’s validity, while an empirical application demonstrates that the previously documented positive effect of the shale gas boom on housing prices is actually driven by trend bias rather than genuine causality. By revealing no true causal impact through credible subgroup analysis, these findings underscore the approach’s critical value in correcting selection bias when standard identifying assumptions are compromised.
This work addresses the lack of a unified configuration governance mechanism in heterogeneous multi-agent systems, which hinders versioned, auditable, and cross-framework consistent management. The authors propose a framework-agnostic reference model for agent configuration governance that normalizes diverse configurations into a canonical configuration graph via semantic projection and enforces uniform governance semantics over this graph. Key innovations include typed and independently versioned configuration items, strict decoupling of configuration from runtime, a lattice-based monotonic influence propagation mechanism, and dependency-aware immutable revisions with provenance tracking. The model is validated across LangGraph, CrewAI, and OpenAI Agents SDK, demonstrating governance-equivalent ACM representations across 27 governance scenarios and 9 propagation cases, while guaranteeing convergence, termination, and a unique fixed point.
This study addresses the lack of systematic evaluation of Programmatic Tool Calling (PTC) versus conventional JSON-based tool invocation under realistic task conditions. Conducting an empirical comparison across 14 prominent language models on the BFCL v4 benchmark, this work provides the first comprehensive validation of PTC’s superiority and robustness across diverse models and scenarios. The methodology exposes tools as typed Python stubs, enabling models to generate executable scripts that support both chained and parallel invocations. Results demonstrate that 11 out of 14 models either maintain or significantly improve performance under PTC, with the GPT-5.6 series achieving a 10.6% gain. Notably, PTC sustains its advantage even in complex settings involving parallel tool use and contextual interference.
This work investigates the dual anchoring phenomenon in vision-language models, wherein model outputs are jointly influenced by textual and visual contexts—even when such contexts are irrelevant or incorrect yet remain plausible in the real world. To this end, we introduce ENTRAP-VL, a novel evaluation framework that establishes the first context-anchoring taxonomy tailored for multimodal settings, explicitly distinguishing between text-driven and vision-driven anchoring mechanisms and incorporating a “scene-inconsistent but world-plausible” realism dimension. Guided by a two-axis design (relevance × realism), we construct a structured probe dataset comprising 1,500 samples across eight scene categories, along with a standardized evaluation protocol. The ENTRAP-VL dataset and associated tools are publicly released to provide the community with a reproducible and extensible foundation for systematic research.