Staff Software Engineer - Compute

Lambda Labs
Bellevue / San Francisco / San Jose2026-08-14Hybrid

About the job

As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.

Responsibilities

- Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.

- Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.

- Guide design of compute platform multi-tenant security model

- Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.

- Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.

- Work with customers on translating vague customer technical requirements into concrete engineering deliverables.

- Set engineering standards and lead design reviews for mission-critical cloud software at scale.

Qualifications

Minimum

- 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.

- Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.

- Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.

- Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.

- Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).

- Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.

Preferred

- Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).

- Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)

- Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).

- Experience with Cloud Service Provider Kubernetes offerings.

- Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).