Institution profile

Zoom Video Communications Inc.

Industry researchnorthamerica · us
Official website
Research library31linked papers
Opportunities0open roles
Selected work

Representative Papers

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Aug 17, 2026

This study addresses the fragmentation of quality estimation and automatic post-editing data for Indian languages by constructing a unified multi-label benchmark comprising 126,000 instances. Through multi-source data integration and stratified sampling, we systematically evaluate large language models and COMET metrics. Results reveal that conflicts between sentence-level and token-level signals serve as a reliable difficulty dimension, while few-shot prompting induces performance degradation and optimal monolingual metrics fail to generalize across language pairs. By establishing standardized evaluation protocols for low-resource translation quality research, this work provides critical empirical evidence to guide future model optimization and metric development.

0 citationsRead paper

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

Jul 16, 2026

This study addresses a critical data degradation issue in reasoning distillation for large language models, where answer-conditioned chain-of-thought generation induces models to favor post-hoc rationalization over genuine forward reasoning. Crucially, this degradation persists even after filtering by posterior correctness. The work systematically uncovers this mechanism for the first time and demonstrates its universality across multiple mainstream large language models through controlled ablation studies, cross-model transfer experiments, prompt ablations, and unsupervised evaluation protocols. Empirical results reveal that training with answer-conditioned chains reduces verifiable reasoning accuracy by up to 27 percentage points on the most challenging competition-level problems, with degradation severity intensifying significantly as problem difficulty increases.

0 citationsRead paper

A Telemetry-Driven Model for Quantifying Upgrade Risk in Durable Workflow Execution

Jul 15, 2026

This work addresses the challenge of state replay failures and silent corruption during version upgrades of long-running workflows, a problem inadequately handled by existing conservative and non-scalable approaches. The authors propose a probabilistic risk assessment model grounded in telemetry data that quantifies upgrade risk using only workflow structural differences and event logs—eliminating the need for sandboxing or shadow execution. Key innovations include the first provably replay-safe migration framework, a coupling graph to model cross-workflow dependencies enabling exact migration strategy partitioning via minimum cut, and an integrated methodology combining static structure comparison, a Bayesian-enhanced Markovian control-flow model, trace-equivalence replay analysis, and fixed-point computation of fault propagation. The system outputs a Workflow Upgrade Risk (WUR) score with confidence intervals, automatically classifying instances into migrate, review, or retain categories, thereby significantly enhancing both safety and efficiency of workflow upgrades.

0 citationsRead paper
Recent publications

Latest Papers

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Aug 17, 2026

This study addresses the fragmentation of quality estimation and automatic post-editing data for Indian languages by constructing a unified multi-label benchmark comprising 126,000 instances. Through multi-source data integration and stratified sampling, we systematically evaluate large language models and COMET metrics. Results reveal that conflicts between sentence-level and token-level signals serve as a reliable difficulty dimension, while few-shot prompting induces performance degradation and optimal monolingual metrics fail to generalize across language pairs. By establishing standardized evaluation protocols for low-resource translation quality research, this work provides critical empirical evidence to guide future model optimization and metric development.

0 citationsRead paper

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

Jul 16, 2026

This study addresses a critical data degradation issue in reasoning distillation for large language models, where answer-conditioned chain-of-thought generation induces models to favor post-hoc rationalization over genuine forward reasoning. Crucially, this degradation persists even after filtering by posterior correctness. The work systematically uncovers this mechanism for the first time and demonstrates its universality across multiple mainstream large language models through controlled ablation studies, cross-model transfer experiments, prompt ablations, and unsupervised evaluation protocols. Empirical results reveal that training with answer-conditioned chains reduces verifiable reasoning accuracy by up to 27 percentage points on the most challenging competition-level problems, with degradation severity intensifying significantly as problem difficulty increases.

0 citationsRead paper

A Telemetry-Driven Model for Quantifying Upgrade Risk in Durable Workflow Execution

Jul 15, 2026

This work addresses the challenge of state replay failures and silent corruption during version upgrades of long-running workflows, a problem inadequately handled by existing conservative and non-scalable approaches. The authors propose a probabilistic risk assessment model grounded in telemetry data that quantifies upgrade risk using only workflow structural differences and event logs—eliminating the need for sandboxing or shadow execution. Key innovations include the first provably replay-safe migration framework, a coupling graph to model cross-workflow dependencies enabling exact migration strategy partitioning via minimum cut, and an integrated methodology combining static structure comparison, a Bayesian-enhanced Markovian control-flow model, trace-equivalence replay analysis, and fixed-point computation of fault propagation. The system outputs a Workflow Upgrade Risk (WUR) score with confidence intervals, automatically classifying instances into migrate, review, or retain categories, thereby significantly enhancing both safety and efficiency of workflow upgrades.

0 citationsRead paper