Institution profile

Oumi AI

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

OpenThoughts-Agent: Data Recipes for Agentic Models

Jun 23, 2026

This work addresses the lack of open, generalizable methodologies for constructing training data for intelligent agents—a key limitation hindering their generalization across diverse tasks. The study presents the first systematic investigation into agent training data construction, introducing an open-source and scalable data recipe. Through multi-source task sampling, diversity optimization, and controlled ablation studies, the authors rigorously analyze how task provenance and data composition influence model performance. A 100K-sample training set built using this approach achieves an average accuracy of 44.8% across seven agent benchmarks, outperforming the strongest existing open-source model by 3.9 percentage points. The method consistently maintains superior performance across varying data scales, demonstrating strong generalization and practical utility.

0 citationsRead paper

Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation

Mar 05, 2026

This work addresses the risk that large language models (LLMs) acting as evaluators may introduce unknown or adversarially exploited biases, thereby compromising the reliability of feedback loops in autonomous AI systems. To mitigate this, the authors propose the Average Bias-Boundedness (A-BB) algorithmic framework—the first approach to formally control LLM evaluator bias with verifiable, bounded guarantees, even when the bias direction is unknown or manipulated adversarially. By integrating statistical hypothesis testing with rank correlation analysis, A-BB is evaluated on the Arena-Hard-Auto benchmark under τ=0.5 and δ=0.01, demonstrating retention of 61%–99% of the original ranking correlation under format- and structure-induced biases, with most configurations exceeding 80%. This balance ensures both safety and practical utility.

0 citationsRead paper

MARVIS: Modality Adaptive Reasoning over VISualizations

Jul 02, 2025

To address poor generalization, domain-specific fine-tuning dependencies, and privacy risks (e.g., PII leakage) in non-traditional modalities (audio, biosignals, tabular data) and long-tailed prediction tasks, this paper proposes MARVIS—a training-free, fine-tuning-free universal cross-modal reasoning framework. MARVIS maps latent representations of arbitrary modalities into visually interpretable images, thereby activating the spatial-structural understanding and fine-grained semantic alignment capabilities of compact vision-language models (VLMs) for zero-shot cross-modal knowledge transfer. Leveraging only a single 3B-parameter VLM, MARVIS achieves an average 16% performance gain over Gemini across diverse long-tailed benchmarks, matching specialized models—while eliminating domain-customized training and exposure of raw sensitive data. To our knowledge, MARVIS is the first framework enabling truly universal multimodal inference with small-scale VLMs under strict privacy-preserving constraints.

0 citationsRead paper
Recent publications

Latest Papers

OpenThoughts-Agent: Data Recipes for Agentic Models

Jun 23, 2026

This work addresses the lack of open, generalizable methodologies for constructing training data for intelligent agents—a key limitation hindering their generalization across diverse tasks. The study presents the first systematic investigation into agent training data construction, introducing an open-source and scalable data recipe. Through multi-source task sampling, diversity optimization, and controlled ablation studies, the authors rigorously analyze how task provenance and data composition influence model performance. A 100K-sample training set built using this approach achieves an average accuracy of 44.8% across seven agent benchmarks, outperforming the strongest existing open-source model by 3.9 percentage points. The method consistently maintains superior performance across varying data scales, demonstrating strong generalization and practical utility.

0 citationsRead paper

Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation

Mar 05, 2026

This work addresses the risk that large language models (LLMs) acting as evaluators may introduce unknown or adversarially exploited biases, thereby compromising the reliability of feedback loops in autonomous AI systems. To mitigate this, the authors propose the Average Bias-Boundedness (A-BB) algorithmic framework—the first approach to formally control LLM evaluator bias with verifiable, bounded guarantees, even when the bias direction is unknown or manipulated adversarially. By integrating statistical hypothesis testing with rank correlation analysis, A-BB is evaluated on the Arena-Hard-Auto benchmark under τ=0.5 and δ=0.01, demonstrating retention of 61%–99% of the original ranking correlation under format- and structure-induced biases, with most configurations exceeding 80%. This balance ensures both safety and practical utility.

0 citationsRead paper

MARVIS: Modality Adaptive Reasoning over VISualizations

Jul 02, 2025

To address poor generalization, domain-specific fine-tuning dependencies, and privacy risks (e.g., PII leakage) in non-traditional modalities (audio, biosignals, tabular data) and long-tailed prediction tasks, this paper proposes MARVIS—a training-free, fine-tuning-free universal cross-modal reasoning framework. MARVIS maps latent representations of arbitrary modalities into visually interpretable images, thereby activating the spatial-structural understanding and fine-grained semantic alignment capabilities of compact vision-language models (VLMs) for zero-shot cross-modal knowledge transfer. Leveraging only a single 3B-parameter VLM, MARVIS achieves an average 16% performance gain over Gemini across diverse long-tailed benchmarks, matching specialized models—while eliminating domain-customized training and exposure of raw sensitive data. To our knowledge, MARVIS is the first framework enabling truly universal multimodal inference with small-scale VLMs under strict privacy-preserving constraints.

0 citationsRead paper