Institution profile

nesa

Research institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments

Aug 08, 2025

To address the high inference cost, poor scalability, and data security risks of large language models (LLMs) in decentralized environments, this paper proposes the first meta-learning-based framework for automated inference acceleration strategy selection. The framework models multi-task historical performance data to learn the adaptivity patterns of various compression, model sharding, and hardware-aware optimization techniques across heterogeneous edge nodes, enabling task- and resource-constrained optimal strategy recommendation. By introducing meta-learning into decentralized inference optimization, it overcomes reliance on manual expertise and inefficient random search. Experiments demonstrate significant improvements over baselines in latency reduction, throughput enhancement, and cross-task generalization—achieving an average 32.7% inference efficiency gain—while maintaining strong practicality and deployability.

0 citationsRead paper

Encrypted Large Model Inference: The Equivariant Encryption Paradigm

Feb 03, 2025

In multi-user settings, deploying large models (e.g., LLMs, diffusion models) on untrusted platforms faces a fundamental privacy–efficiency trade-off. Method: This paper introduces Equivariant Encryption (EE), a novel paradigm grounded in the theory of function equivariance under group actions; EE selectively obfuscates critical layer representations while enabling exact ciphertext computation for linear and prescribed nonlinear operations—bypassing the prohibitive overhead of fully homomorphic encryption (FHE). EE is architecture-agnostic, supporting CNNs, Transformers, and others, and integrates seamlessly into standard inference pipelines. Contribution/Results: Experiments demonstrate end-to-end privacy preservation in decentralized scenarios: inputs, intermediate activations, and outputs remain confidential throughout inference. Accuracy is preserved losslessly; throughput approaches plaintext-level performance, with latency overhead <0.5%. EE significantly outperforms secure multi-party computation (SMPC) and FHE in both efficiency and practicality.

0 citationsRead paper
Recent publications

Latest Papers

Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments

Aug 08, 2025

To address the high inference cost, poor scalability, and data security risks of large language models (LLMs) in decentralized environments, this paper proposes the first meta-learning-based framework for automated inference acceleration strategy selection. The framework models multi-task historical performance data to learn the adaptivity patterns of various compression, model sharding, and hardware-aware optimization techniques across heterogeneous edge nodes, enabling task- and resource-constrained optimal strategy recommendation. By introducing meta-learning into decentralized inference optimization, it overcomes reliance on manual expertise and inefficient random search. Experiments demonstrate significant improvements over baselines in latency reduction, throughput enhancement, and cross-task generalization—achieving an average 32.7% inference efficiency gain—while maintaining strong practicality and deployability.

0 citationsRead paper

Encrypted Large Model Inference: The Equivariant Encryption Paradigm

Feb 03, 2025

In multi-user settings, deploying large models (e.g., LLMs, diffusion models) on untrusted platforms faces a fundamental privacy–efficiency trade-off. Method: This paper introduces Equivariant Encryption (EE), a novel paradigm grounded in the theory of function equivariance under group actions; EE selectively obfuscates critical layer representations while enabling exact ciphertext computation for linear and prescribed nonlinear operations—bypassing the prohibitive overhead of fully homomorphic encryption (FHE). EE is architecture-agnostic, supporting CNNs, Transformers, and others, and integrates seamlessly into standard inference pipelines. Contribution/Results: Experiments demonstrate end-to-end privacy preservation in decentralized scenarios: inputs, intermediate activations, and outputs remain confidential throughout inference. Accuracy is preserved losslessly; throughput approaches plaintext-level performance, with latency overhead <0.5%. EE significantly outperforms secure multi-party computation (SMPC) and FHE in both efficiency and practicality.

0 citationsRead paper