Toward Super Agent System with Hybrid AI Routers
To address high latency, elevated operational costs, and privacy risks in deploying large-scale super-agent systems in real-world settings, this paper proposes a cloud-edge collaborative hybrid AI router architecture. The architecture employs lightweight intent recognition and multi-level task routing to dynamically dispatch user queries either to specialized agents or to automatically generated workflows, while adaptively orchestrating edge-deployed lightweight models and cloud-based large language models based on task complexity. Its key innovations include heterogeneous model cooperative inference, joint edge-cloud optimization, and multimodal model lightweight adaptation. Experimental evaluation demonstrates that the system significantly reduces on-device response latency (average reduction of 42%) and cloud service expenditure (cost reduction of 38%) across edge devices such as smartphones and robots. By simultaneously achieving low latency, cost efficiency, and enhanced data privacy, the proposed architecture delivers a scalable, production-ready systems-level solution for super-agent deployment.