Institution profile

Transmute AI Lab

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models

Jan 02, 2026Trans. Mach. Learn. Res.

This work addresses the susceptibility of vision-language models to hallucination during generation, which undermines their practical reliability. To mitigate this issue without requiring additional training, the authors propose a novel framework that constructs diverse hallucinatory variants by selectively removing critical textual tokens and then fuses multi-source hallucination signals through generalized contrastive decoding. This approach overcomes the limitations of existing methods that rely on overly narrow assumptions about hallucination origins. Extensive experiments demonstrate consistent performance gains across six benchmark datasets and three prominent vision-language models, achieving a 20% improvement in CHAI-R score and outperforming current state-of-the-art training-free hallucination suppression techniques.

0 citationsRead paper

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation

Oct 08, 2025

Existing LLM reinforcement learning methods (e.g., GRPO) suffer from insufficient exploration under complex prompts and inefficient utilization of sparse rewards, primarily due to fixed rollout allocation and context-agnostic policies. To address these limitations, we propose XRPO—a novel framework featuring three key innovations: (1) uncertainty-driven adaptive rollout allocation, dynamically increasing exploration for low-reward prompts; (2) in-context seed guidance to activate high-quality response generation; and (3) a sequence-likelihood-aware group relative advantage function coupled with novelty-aware advantage sharpening, enhancing detection of sparse correct responses. Evaluated on mathematical reasoning and programming benchmarks, XRPO significantly outperforms baselines including GRPO and GSPO—achieving up to +4% absolute gain in pass@1 and +6% in cons@32—while accelerating training convergence by 2.7×.

0 citationsRead paper
Recent publications

Latest Papers

CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models

Jan 02, 2026Trans. Mach. Learn. Res.

This work addresses the susceptibility of vision-language models to hallucination during generation, which undermines their practical reliability. To mitigate this issue without requiring additional training, the authors propose a novel framework that constructs diverse hallucinatory variants by selectively removing critical textual tokens and then fuses multi-source hallucination signals through generalized contrastive decoding. This approach overcomes the limitations of existing methods that rely on overly narrow assumptions about hallucination origins. Extensive experiments demonstrate consistent performance gains across six benchmark datasets and three prominent vision-language models, achieving a 20% improvement in CHAI-R score and outperforming current state-of-the-art training-free hallucination suppression techniques.

0 citationsRead paper

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation

Oct 08, 2025

Existing LLM reinforcement learning methods (e.g., GRPO) suffer from insufficient exploration under complex prompts and inefficient utilization of sparse rewards, primarily due to fixed rollout allocation and context-agnostic policies. To address these limitations, we propose XRPO—a novel framework featuring three key innovations: (1) uncertainty-driven adaptive rollout allocation, dynamically increasing exploration for low-reward prompts; (2) in-context seed guidance to activate high-quality response generation; and (3) a sequence-likelihood-aware group relative advantage function coupled with novelty-aware advantage sharpening, enhancing detection of sparse correct responses. Evaluated on mathematical reasoning and programming benchmarks, XRPO significantly outperforms baselines including GRPO and GSPO—achieving up to +4% absolute gain in pass@1 and +6% in cons@32—while accelerating training convergence by 2.7×.

0 citationsRead paper