Engineering Manager, Fleet Engineering

Lambda Labs
San Francisco, CA, USA / San Jose, CA, USA / Bellevue, WA, USA2026-08-22Hybrid

About the job

Fleet Engineering owns the full lifecycle of Lambda's production systems infrastructure — new product introduction, deployment, operation, and reliability of our GPU fleet. We enable the building and running of that infrastructure with speed, ease, and quality. The work is highly cross-functional, carries executive visibility, and has a direct impact on Lambda and our customers. Fleet Engineering is at the forefront of delivering on-time, high-quality GPU capacity while driving efficiency at scale. We are hiring multiple Engineering Managers for the following teams: Fleet Reliability, HPC Deployments, Fleet Foundation, Fleet Orchestration / Automation.

Responsibilities

- Lead and grow a distributed team of top-talent engineers responsible for the deployment and operation of production systems infrastructure.

- Work cross-functionally to deliver projects and deployments on time, ensuring alignment across stakeholders.

- Identify opportunities for efficiency gains in the tools, processes, and automation that teams across the organization rely on day to day.

- Give stakeholders clear visibility into project progress, risks, and outcomes.

- Participate in qualification efforts for new technologies entering our production deployments.

- Drive outcomes by managing staff allocation, project priorities, deadlines, and deliverables.

- Hold regular 1:1s, give constructive feedback, and support career development for your team.

- Contribute to reliability through participation in our Incident Management and Review programs.

Qualifications

Minimum

- Have 3+ years leading or managing engineers, in AI/ML infrastructure or another large-scale compute environment.

- Have owned production systems with real SLAs, and can balance keeping things running against long-term, high-impact work — paying down toil and technical debt along the way.

- Work confidently in Linux and can debug across the OS, hardware, and networking layers.

- Can lead technical design on medium-to-large efforts: take an ambiguous problem, write the doc, drive alignment across teams, and ship.

- Work well under deadlines and structured project plans, and can tactfully negotiate changes to timelines when reality demands it.

- Collaborate effectively with peer engineering managers on efforts that cut across deployment and operations.

- Build high-performing teams deliberately — through hiring, upskilling, planned skills redundancy, performance management, and clear expectations.

- Have excellent problem-solving and troubleshooting instincts.

- Are excited about working at the intersection of hardware, software, and physical datacenter builds.

- Leave systems, and the teammates around you, better than you found them.

Preferred

- Linux systems administration, TCP/IP networking, automation, and scripting.

- Bare metal provisioning and lifecycle management — PXE, Redfish, IPMI, BMC, DHCP, DNS.

- Strong coding ability in at least one language, plus comfort with APIs, distributed systems, and automation pipelines.

- The technologies underpinning our cloud business: GPU acceleration, virtualization, cloud computing.

- Datacenter physical infrastructure: racks, switches, InfiniBand fabric, power domains.

- Network source-of-truth or DCIM tooling (NetBox or similar), and data quality practice at scale.

- Building Linux distributions, or managing OS customization and imaging.

- Incorporating AI-assisted development tools into engineering workflows — code generation, debugging, test development, documentation.

- Customer awareness, empathy, and diplomacy.

- Bachelor's degree or equivalent experience in a technical field.