Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
本文提出一个框架,用于分析AI代理跨组织边界交互时出现的风险,并探讨了三种部署层级下的风险因素、失败模式及控制措施。
本文提出一个框架,用于分析AI代理跨组织边界交互时出现的风险,并探讨了三种部署层级下的风险因素、失败模式及控制措施。
This work addresses the limited cost-effectiveness of existing routing methods that merely assign simple tasks to small models without enhancing their capabilities. To overcome this, the authors propose a multi-cycle adaptation mechanism operating at the granularity of single inference calls. The approach leverages a teacher model to generate verification demonstrations from the small model’s failures, integrating skill distillation and LoRA fine-tuning to continuously improve its competence. Joint optimization is performed over a dynamic skill library, task-specific adapters, and a cost-calibrated routing policy, complemented by a verifier-supported fallback mechanism. Experiments show that Qwen2.5-Coder-1.5B achieves a pass rate increase from 28.7% to 49.7% on HumanEval+MBPP; the deployment strategy attains 88.3% of peak performance at only 60.8% of the cost; and Qwen3.5-2B matches the performance of an unadapted 4B model on TAU-2.
本文提出一个框架,用于分析AI代理跨组织边界交互时出现的风险,并探讨了三种部署层级下的风险因素、失败模式及控制措施。
This work addresses the limited cost-effectiveness of existing routing methods that merely assign simple tasks to small models without enhancing their capabilities. To overcome this, the authors propose a multi-cycle adaptation mechanism operating at the granularity of single inference calls. The approach leverages a teacher model to generate verification demonstrations from the small model’s failures, integrating skill distillation and LoRA fine-tuning to continuously improve its competence. Joint optimization is performed over a dynamic skill library, task-specific adapters, and a cost-calibrated routing policy, complemented by a verifier-supported fallback mechanism. Experiments show that Qwen2.5-Coder-1.5B achieves a pass rate increase from 28.7% to 49.7% on HumanEval+MBPP; the deployment strategy attains 88.3% of peak performance at only 60.8% of the cost; and Qwen3.5-2B matches the performance of an unadapted 4B model on TAU-2.