An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决云LLM服务中的隐私风险,本文提出Odin系统,通过共同设计密文打包和模型执行优化CKKS-based LLM推理,显著减少计算时间和内存消耗。
📝 Abstract
Cloud LLM services typically require users to send prompts to a model provider, creating a privacy risk. Fully homomorphic encryption (FHE) lets a server perform inference without decrypting the input, but representing data as ciphertexts adds storage and computational overhead. In CKKS-based LLM inference, the packing scheme maps logical tensors to ciphertexts and slots. It therefore determines the ciphertext count and the homomorphic cost of linear layers, and it constrains how data pass between linear layers, attention, and nonlinear computation. As models and sequences grow, inefficient layouts accumulate encoding, compute, and layout-conversion overhead. We present Odin, an FHE inference system that co-designs ciphertext packing and model execution for Llama. Starting from a THOR-style baseline whose bottleneck is weight encoding, Odin uses a feature-major cross-layer layout to unify residual connections and layer interfaces, and builds transient intra-operator layouts for linear projections and attention. This reduces redundant plaintext encoding of weights in wide projections. Within attention, QK^T produces scores that Softmax can consume directly, and PV consumes the resulting probabilities, avoiding intermediate repacking. For nonlinear ops, we use minimax polynomial approximation with input-range control and joint error allocation guided by model quality, reducing polynomial degree and multiplicative depth. To our knowledge, Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3. With Llama-3-8B weights and a 128-token input, Odin evaluates all 32 Transformer layers on a single NVIDIA H100 80 GB GPU. Server-side end-to-end FHE evaluation takes 366.4 s and 58.9 GiB peak device memory. Under the same model, input, CKKS parameters, and hardware, THOR takes 1651.9 s, a 4.51x speedup.
Problem

Research questions and friction points this paper is trying to address.

Fully Homomorphic Encryption
Privacy-Preserving Inference
Ciphertext Packing
LLM
CKKS
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fully Homomorphic Encryption
Ciphertext Packing
Minimax Polynomial Approximation
Attention Mechanism Optimization
🔎 Similar Papers
2024-10-03arXiv.orgCitations: 0
Y
Yuhang Fan
School of Cyber Science and Technology, Shandong University
Y
Yusi Chen
School of Cyber Science and Technology, Shandong University
K
Kanyu Ye
School of Cyber Science and Technology, Shandong University
Zhuoran Ji
Zhuoran Ji
Shandong University
GPUcompilerhigh performance computing