🤖 AI Summary
本文提出NACRE,一种RISC-V硬件-软件协同设计,通过分离主机管理资源与访问保护状态的权限,实现原生保密容器,解决了现有系统对Linux容器保护不足的问题。
📝 Abstract
Linux containers achieve high density and fast lifecycle operations by sharing the host kernel, but this design also lets a compromised host inspect or modify container state. Existing
confidential-computing systems protect an enclave address space or an entire guest operating system, while recent container-granularity systems still add a separate protection context.
These abstractions do not make a dynamic group of host-managed Linux processes the architectural protection unit.
This paper presents NACRE, a RISC-V hardware-software co-design for native confidential containers. Its key insight is to separate the host's authority to manage resources from its
authority to access or commit protected state. Hardware-recognized container identities direct protected traps to an isolated S-mode agent, while an M-mode monitor commits security-
sensitive identity, mapping, and page transitions. The agent delegates services to host Linux without changing satp; services that neither access private bytes nor modify protected
state also avoid M-mode. We prototype NACRE by extending QEMU, OpenSBI, Linux, a trusted agent, and runc. The prototype implements the single-container private-memory substrate and
covered launch, fault, fork/COW, user-access, and teardown paths. Across five lmbench syscall and pipe metrics, the three-run means remain within 3.5% of the runc-origin baseline. With
the eight nginx object-size means weighted equally, aggregate throughput is 1.9% lower.