Institution profile

Gradient

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

Aug 10, 2026

This work addresses the limited cost-effectiveness of existing routing methods that merely assign simple tasks to small models without enhancing their capabilities. To overcome this, the authors propose a multi-cycle adaptation mechanism operating at the granularity of single inference calls. The approach leverages a teacher model to generate verification demonstrations from the small model’s failures, integrating skill distillation and LoRA fine-tuning to continuously improve its competence. Joint optimization is performed over a dynamic skill library, task-specific adapters, and a cost-calibrated routing policy, complemented by a verifier-supported fallback mechanism. Experiments show that Qwen2.5-Coder-1.5B achieves a pass rate increase from 28.7% to 49.7% on HumanEval+MBPP; the deployment strategy attains 88.3% of peak performance at only 60.8% of the cost; and Qwen3.5-2B matches the performance of an unadapted 4B model on TAU-2.

0 citationsRead paper
Recent publications

Latest Papers

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

Aug 10, 2026

This work addresses the limited cost-effectiveness of existing routing methods that merely assign simple tasks to small models without enhancing their capabilities. To overcome this, the authors propose a multi-cycle adaptation mechanism operating at the granularity of single inference calls. The approach leverages a teacher model to generate verification demonstrations from the small model’s failures, integrating skill distillation and LoRA fine-tuning to continuously improve its competence. Joint optimization is performed over a dynamic skill library, task-specific adapters, and a cost-calibrated routing policy, complemented by a verifier-supported fallback mechanism. Experiments show that Qwen2.5-Coder-1.5B achieves a pass rate increase from 28.7% to 49.7% on HumanEval+MBPP; the deployment strategy attains 88.3% of peak performance at only 60.8% of the cost; and Qwen3.5-2B matches the performance of an unadapted 4B model on TAU-2.

0 citationsRead paper