State Without a Landlord: An Architecture Proposal for Peer-to-Peer Replication of Durable Workflow State
本文提出了一种基于对等复制的持久工作流状态架构,通过认证的追加日志和显式证书来解决现有框架中状态权威性和信任问题。
本文提出了一种基于对等复制的持久工作流状态架构,通过认证的追加日志和显式证书来解决现有框架中状态权威性和信任问题。
为了解决视觉语言模型中推测解码的效率问题,提出了一种名为GLANCE的一次性块草稿方法,该方法通过融合视觉-语言状态并在一个前向传递中填充整个块来提高解码速度。
This study addresses the fragmentation of quality estimation and automatic post-editing data for Indian languages by constructing a unified multi-label benchmark comprising 126,000 instances. Through multi-source data integration and stratified sampling, we systematically evaluate large language models and COMET metrics. Results reveal that conflicts between sentence-level and token-level signals serve as a reliable difficulty dimension, while few-shot prompting induces performance degradation and optimal monolingual metrics fail to generalize across language pairs. By establishing standardized evaluation protocols for low-resource translation quality research, this work provides critical empirical evidence to guide future model optimization and metric development.
This study addresses a critical data degradation issue in reasoning distillation for large language models, where answer-conditioned chain-of-thought generation induces models to favor post-hoc rationalization over genuine forward reasoning. Crucially, this degradation persists even after filtering by posterior correctness. The work systematically uncovers this mechanism for the first time and demonstrates its universality across multiple mainstream large language models through controlled ablation studies, cross-model transfer experiments, prompt ablations, and unsupervised evaluation protocols. Empirical results reveal that training with answer-conditioned chains reduces verifiable reasoning accuracy by up to 27 percentage points on the most challenging competition-level problems, with degradation severity intensifying significantly as problem difficulty increases.
This work addresses the challenge of state replay failures and silent corruption during version upgrades of long-running workflows, a problem inadequately handled by existing conservative and non-scalable approaches. The authors propose a probabilistic risk assessment model grounded in telemetry data that quantifies upgrade risk using only workflow structural differences and event logs—eliminating the need for sandboxing or shadow execution. Key innovations include the first provably replay-safe migration framework, a coupling graph to model cross-workflow dependencies enabling exact migration strategy partitioning via minimum cut, and an integrated methodology combining static structure comparison, a Bayesian-enhanced Markovian control-flow model, trace-equivalence replay analysis, and fixed-point computation of fault propagation. The system outputs a Workflow Upgrade Risk (WUR) score with confidence intervals, automatically classifying instances into migrate, review, or retain categories, thereby significantly enhancing both safety and efficiency of workflow upgrades.
本文提出了一种基于对等复制的持久工作流状态架构,通过认证的追加日志和显式证书来解决现有框架中状态权威性和信任问题。
为了解决视觉语言模型中推测解码的效率问题,提出了一种名为GLANCE的一次性块草稿方法,该方法通过融合视觉-语言状态并在一个前向传递中填充整个块来提高解码速度。
This study addresses the fragmentation of quality estimation and automatic post-editing data for Indian languages by constructing a unified multi-label benchmark comprising 126,000 instances. Through multi-source data integration and stratified sampling, we systematically evaluate large language models and COMET metrics. Results reveal that conflicts between sentence-level and token-level signals serve as a reliable difficulty dimension, while few-shot prompting induces performance degradation and optimal monolingual metrics fail to generalize across language pairs. By establishing standardized evaluation protocols for low-resource translation quality research, this work provides critical empirical evidence to guide future model optimization and metric development.
This study addresses a critical data degradation issue in reasoning distillation for large language models, where answer-conditioned chain-of-thought generation induces models to favor post-hoc rationalization over genuine forward reasoning. Crucially, this degradation persists even after filtering by posterior correctness. The work systematically uncovers this mechanism for the first time and demonstrates its universality across multiple mainstream large language models through controlled ablation studies, cross-model transfer experiments, prompt ablations, and unsupervised evaluation protocols. Empirical results reveal that training with answer-conditioned chains reduces verifiable reasoning accuracy by up to 27 percentage points on the most challenging competition-level problems, with degradation severity intensifying significantly as problem difficulty increases.
This work addresses the challenge of state replay failures and silent corruption during version upgrades of long-running workflows, a problem inadequately handled by existing conservative and non-scalable approaches. The authors propose a probabilistic risk assessment model grounded in telemetry data that quantifies upgrade risk using only workflow structural differences and event logs—eliminating the need for sandboxing or shadow execution. Key innovations include the first provably replay-safe migration framework, a coupling graph to model cross-workflow dependencies enabling exact migration strategy partitioning via minimum cut, and an integrated methodology combining static structure comparison, a Bayesian-enhanced Markovian control-flow model, trace-equivalence replay analysis, and fixed-point computation of fault propagation. The system outputs a Workflow Upgrade Risk (WUR) score with confidence intervals, automatically classifying instances into migrate, review, or retain categories, thereby significantly enhancing both safety and efficiency of workflow upgrades.