Institution profile

Reka AI

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities7open roles
Selected work

Representative Papers

Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models

May 31, 2026

This work addresses the limited multilingual reasoning capabilities of large language models on low-resource and unseen languages, as well as the absence of effective methods that operate without labeled or parallel data. The authors propose an unsupervised reinforcement learning framework that enhances reasoning performance by enforcing cross-lingual self-consistency—requiring the model to produce consistent answers to semantically equivalent questions across languages—without relying on gold labels or multilingual alignment data. This approach substantially improves generalization to both unseen languages and out-of-distribution tasks, achieving an average accuracy gain of 21.7% across the ten languages in the MGSM benchmark, with an 18.2% improvement specifically on unseen languages, and up to a 6.2% increase on three out-of-distribution evaluation benchmarks.

0 citationsRead paper

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training

Jul 29, 2025

This study addresses the challenge of evaluating data source quality and allocating resources in domain-adaptive pretraining—specifically, how to reliably and cost-effectively compare the long-term utility of heterogeneous data sources (e.g., synthetic, web-scraped, and user-generated data) across varying compute scales. We propose a scaling-law-based framework for data source utility assessment, which fits performance scaling curves via multi-round annealing training at diverse compute budgets. This approach overcomes the instability of point-wise estimates when ranking sources across scales. The framework enables consistent, cross-source and cross-budget utility comparison. Evaluated on a 7B-parameter model, it accurately predicts long-horizon performance gains, significantly improving specialization in medicine and mathematics. Moreover, it informs cost-optimal data curation and training strategies.

0 citationsRead paper

Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation

May 30, 2025

This study investigates the impact of English data inclusion in continual pretraining (CPT) on multilingual large language models’ adaptation to new target languages. It reveals that while mixing English data does not degrade validation perplexity, it significantly enhances in-context learning (ICL) and generalization in the target language; conversely, monolingual CPT often triggers catastrophic forgetting. Method: To diagnose this phenomenon, the authors introduce the first language-agnostic ICL evaluation benchmark and propose a novel English-free CPT paradigm integrating curriculum learning with weight exponential moving average (EMA). Contribution/Results: Experiments demonstrate that the proposed method effectively mitigates forgetting, stabilizes downstream task performance, and suppresses parameter drift. It provides a reproducible, lightweight technical pathway for efficient adaptation of multilingual LMs to low-resource languages.

0 citationsRead paper

WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging

Feb 25, 2025

Existing multiple-choice benchmarks suffer from insufficient difficulty, limiting their ability to discriminate models’ true capabilities. To address this, we propose WiCkeD—a method that systematically injects semantic consistency and reasoning robustness challenges by automatically replacing any correct option with “None of the Above” (NOTA). This perturbation is agnostic to benchmark design and seamlessly integrates with arbitrary multiple-choice evaluation suites, enabling scalable assessment of open large language models. Evaluated across six mainstream benchmarks, WiCkeD induces an average accuracy drop of 12.1 points across 18 models; even chain-of-thought (CoT)-enhanced models exhibit significant degradation, exposing their sensitivity to implicit logical consistency. To our knowledge, this is the first work to formalize NOTA as a controllable difficulty-augmentation mechanism, offering fine-grained capability analysis beyond aggregate accuracy metrics. All code and data are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models

May 31, 2026

This work addresses the limited multilingual reasoning capabilities of large language models on low-resource and unseen languages, as well as the absence of effective methods that operate without labeled or parallel data. The authors propose an unsupervised reinforcement learning framework that enhances reasoning performance by enforcing cross-lingual self-consistency—requiring the model to produce consistent answers to semantically equivalent questions across languages—without relying on gold labels or multilingual alignment data. This approach substantially improves generalization to both unseen languages and out-of-distribution tasks, achieving an average accuracy gain of 21.7% across the ten languages in the MGSM benchmark, with an 18.2% improvement specifically on unseen languages, and up to a 6.2% increase on three out-of-distribution evaluation benchmarks.

0 citationsRead paper

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training

Jul 29, 2025

This study addresses the challenge of evaluating data source quality and allocating resources in domain-adaptive pretraining—specifically, how to reliably and cost-effectively compare the long-term utility of heterogeneous data sources (e.g., synthetic, web-scraped, and user-generated data) across varying compute scales. We propose a scaling-law-based framework for data source utility assessment, which fits performance scaling curves via multi-round annealing training at diverse compute budgets. This approach overcomes the instability of point-wise estimates when ranking sources across scales. The framework enables consistent, cross-source and cross-budget utility comparison. Evaluated on a 7B-parameter model, it accurately predicts long-horizon performance gains, significantly improving specialization in medicine and mathematics. Moreover, it informs cost-optimal data curation and training strategies.

0 citationsRead paper

Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation

May 30, 2025

This study investigates the impact of English data inclusion in continual pretraining (CPT) on multilingual large language models’ adaptation to new target languages. It reveals that while mixing English data does not degrade validation perplexity, it significantly enhances in-context learning (ICL) and generalization in the target language; conversely, monolingual CPT often triggers catastrophic forgetting. Method: To diagnose this phenomenon, the authors introduce the first language-agnostic ICL evaluation benchmark and propose a novel English-free CPT paradigm integrating curriculum learning with weight exponential moving average (EMA). Contribution/Results: Experiments demonstrate that the proposed method effectively mitigates forgetting, stabilizes downstream task performance, and suppresses parameter drift. It provides a reproducible, lightweight technical pathway for efficient adaptation of multilingual LMs to low-resource languages.

0 citationsRead paper

WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging

Feb 25, 2025

Existing multiple-choice benchmarks suffer from insufficient difficulty, limiting their ability to discriminate models’ true capabilities. To address this, we propose WiCkeD—a method that systematically injects semantic consistency and reasoning robustness challenges by automatically replacing any correct option with “None of the Above” (NOTA). This perturbation is agnostic to benchmark design and seamlessly integrates with arbitrary multiple-choice evaluation suites, enabling scalable assessment of open large language models. Evaluated across six mainstream benchmarks, WiCkeD induces an average accuracy drop of 12.1 points across 18 models; even chain-of-thought (CoT)-enhanced models exhibit significant degradation, exposing their sensitivity to implicit logical consistency. To our knowledge, this is the first work to formalize NOTA as a controllable difficulty-augmentation mechanism, offering fine-grained capability analysis beyond aggregate accuracy metrics. All code and data are publicly released.

0 citationsRead paper