About the job
Fleet Engineering owns the full lifecycle of Lambda's production systems infrastructure — new product introduction, deployment, operation, and reliability of our GPU fleet. We enable the building and running of that infrastructure with speed, ease, and quality. The work is highly cross-functional, carries executive visibility, and has a direct impact on Lambda and our customers. Fleet Engineering is at the forefront of delivering on-time, high-quality GPU capacity while driving efficiency at scale. We are hiring multiple Engineering Managers for the following teams: Fleet Reliability, HPC Deployments, Fleet Foundation, Fleet Orchestration / Automation.
Responsibilities
- Lead and grow a distributed team of top-talent engineers responsible for the deployment and operation of production systems infrastructure.
- Work cross-functionally to deliver projects and deployments on time, ensuring alignment across stakeholders.
- Identify opportunities for efficiency gains in the tools, processes, and automation that teams across the organization rely on day to day.
- Give stakeholders clear visibility into project progress, risks, and outcomes.
- Participate in qualification efforts for new technologies entering our production deployments.
- Drive outcomes by managing staff allocation, project priorities, deadlines, and deliverables.
- Hold regular 1:1s, give constructive feedback, and support career development for your team.
- Contribute to reliability through participation in our Incident Management and Review programs.
Qualifications
Minimum
- Have 3+ years leading or managing engineers, in AI/ML infrastructure or another large-scale compute environment.
- Have owned production systems with real SLAs, and can balance keeping things running against long-term, high-impact work — paying down toil and technical debt along the way.
- Work confidently in Linux and can debug across the OS, hardware, and networking layers.
- Can lead technical design on medium-to-large efforts: take an ambiguous problem, write the doc, drive alignment across teams, and ship.
- Work well under deadlines and structured project plans, and can tactfully negotiate changes to timelines when reality demands it.
- Collaborate effectively with peer engineering managers on efforts that cut across deployment and operations.
- Build high-performing teams deliberately — through hiring, upskilling, planned skills redundancy, performance management, and clear expectations.
- Have excellent problem-solving and troubleshooting instincts.
- Are excited about working at the intersection of hardware, software, and physical datacenter builds.
- Leave systems, and the teammates around you, better than you found them.
Preferred
- Linux systems administration, TCP/IP networking, automation, and scripting.
- Bare metal provisioning and lifecycle management — PXE, Redfish, IPMI, BMC, DHCP, DNS.
- Strong coding ability in at least one language, plus comfort with APIs, distributed systems, and automation pipelines.
- The technologies underpinning our cloud business: GPU acceleration, virtualization, cloud computing.
- Datacenter physical infrastructure: racks, switches, InfiniBand fabric, power domains.
- Network source-of-truth or DCIM tooling (NetBox or similar), and data quality practice at scale.
- Building Linux distributions, or managing OS customization and imaging.
- Incorporating AI-assisted development tools into engineering workflows — code generation, debugging, test development, documentation.
- Customer awareness, empathy, and diplomacy.
- Bachelor's degree or equivalent experience in a technical field.