About the job
As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.
Responsibilities
- Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.
- Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
- Guide design of compute platform multi-tenant security model
- Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
- Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.
- Work with customers on translating vague customer technical requirements into concrete engineering deliverables.
- Set engineering standards and lead design reviews for mission-critical cloud software at scale.
Qualifications
Minimum
- 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
- Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
- Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
- Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
- Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
- Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.
Preferred
- Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
- Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
- Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).
- Experience with Cloud Service Provider Kubernetes offerings.
- Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).