Institution profile

TensorOpera Inc

Industry researchnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Toward Super Agent System with Hybrid AI Routers

Apr 11, 2025

To address high latency, elevated operational costs, and privacy risks in deploying large-scale super-agent systems in real-world settings, this paper proposes a cloud-edge collaborative hybrid AI router architecture. The architecture employs lightweight intent recognition and multi-level task routing to dynamically dispatch user queries either to specialized agents or to automatically generated workflows, while adaptively orchestrating edge-deployed lightweight models and cloud-based large language models based on task complexity. Its key innovations include heterogeneous model cooperative inference, joint edge-cloud optimization, and multimodal model lightweight adaptation. Experimental evaluation demonstrates that the system significantly reduces on-device response latency (average reduction of 42%) and cloud service expenditure (cost reduction of 38%) across edge devices such as smartphones and robots. By simultaneously achieving low latency, cost efficiency, and enhanced data privacy, the proposed architecture delivers a scalable, production-ready systems-level solution for super-agent deployment.

0 citationsRead paper
Recent publications

Latest Papers

Toward Super Agent System with Hybrid AI Routers

Apr 11, 2025

To address high latency, elevated operational costs, and privacy risks in deploying large-scale super-agent systems in real-world settings, this paper proposes a cloud-edge collaborative hybrid AI router architecture. The architecture employs lightweight intent recognition and multi-level task routing to dynamically dispatch user queries either to specialized agents or to automatically generated workflows, while adaptively orchestrating edge-deployed lightweight models and cloud-based large language models based on task complexity. Its key innovations include heterogeneous model cooperative inference, joint edge-cloud optimization, and multimodal model lightweight adaptation. Experimental evaluation demonstrates that the system significantly reduces on-device response latency (average reduction of 42%) and cloud service expenditure (cost reduction of 38%) across edge devices such as smartphones and robots. By simultaneously achieving low latency, cost efficiency, and enhanced data privacy, the proposed architecture delivers a scalable, production-ready systems-level solution for super-agent deployment.

0 citationsRead paper