BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
为解决多语言模型在小语种上表现不佳的问题,通过针对102种语言进行单语微调及更复杂的语言适应策略,提高了自动语音识别的准确性。
为解决多语言模型在小语种上表现不佳的问题,通过针对102种语言进行单语微调及更复杂的语言适应策略,提高了自动语音识别的准确性。
研究通过测试严格单语模型,发现无需联合训练也能学习可对齐的表示,揭示了跨语言对齐可从语言结构中自然产生。
研究通过多语言自我博弈方法,量化了大型语言模型在不同语言间技能表现的不一致性问题。
本文探讨了多语言模型评价中的公平比较问题,通过控制变量实验揭示常用指标存在的偏差,并提出基于语义等价序列的句子级负对数似然作为更有效的跨语言比较方法。
This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.
为解决多语言模型在小语种上表现不佳的问题,通过针对102种语言进行单语微调及更复杂的语言适应策略,提高了自动语音识别的准确性。
研究通过测试严格单语模型,发现无需联合训练也能学习可对齐的表示,揭示了跨语言对齐可从语言结构中自然产生。
研究通过多语言自我博弈方法,量化了大型语言模型在不同语言间技能表现的不一致性问题。
本文探讨了多语言模型评价中的公平比较问题,通过控制变量实验揭示常用指标存在的偏差,并提出基于语义等价序列的句子级负对数似然作为更有效的跨语言比较方法。
This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.