About the job
Crusoe is the world's first vertically integrated, sustainable AI cloud. We build and operate GPU infrastructure powered by clean energy, from data center design through IaaS products to managed inference at scale, enabling AI-native companies to run demanding workloads without compromising on sustainability or reliability. Crusoe Cloud is 1,400 people and growing. The TPM frameworks are still being built, which means there is a real opportunity to shape how the function operates rather than inherit how it already works.
We are hiring a Staff TPM who will own deployment programs. You will be helping us build out new sites, capacity expansion, or help deploy modular (Spark) data centers. Our TPMs define and drive the entire deployment engagement model from chip vendor engagement through first customer cluster delivery.
Responsibilities
- Own the infrastructure deployment for new sites or site expansions end-to-end: chip vendor and OEM dependencies, architecture updates, cloud foundations work, commissioning gate framework definition, and first customer cluster delivery as the success metric.
- Lead Deployment Phase 0 on the Cloud TPM side: define firmware version targets and DOCA targets before kickoff, set commissioning gate criteria, and define the DRI matrix.
- Manage compounding cross-SKU dependencies where active production programs (e.g. B200/GB300/VR) are running in parallel with capacity expansion projects, and prevent them from competing for the same Engineering pool without a plan.
- Owns real-time execution dashboards; delivers crisp, data-driven executive updates that surface decision elements without requiring follow-up
- Governs cross-organizational dependencies without waiting for escalation authority
- Coach more junior TPMs on technical depth, risk identification, and executive communication.
- Actively drives AI tool integration across their programs; identifies where AI materially improves program tracking, risk detection, and executive communication
Qualifications
Minimum
- Deep, working fluency with GPU architecture across SKU generations, firmware lifecycle (DOCA, driver stacks, BIOS/BMC), compute orchestration, SDN, storage, networking (leaf-spine topology, ZTP, fabric commissioning), and monitoring/observability
- Direct hardware partner engagement: personal ownership of NVIDIA or OEM certification and validation timelines, not coordination feeding into someone else's relationship.
- Active daily use of AI tools to drive program-level outcomes: risk detection, dependency mapping, data analysis, and executive communication, not just personal productivity.
- 10+ years as a Technical Program Manager with a track record of owning infrastructure deployment programs end-to-end at a hyperscaler, GPU cloud provider, or AI infrastructure company, ideally with direct experience in large-scale capacity expansion projects.
- Proven ability to define deployment engagement models from scratch, not just operate within existing frameworks, and make them stick across engineering organizations that didn't ask for them.
- Track record of driving cross-organizational alignment at VP/SVP level without formal authority, including building durable alignment on programs that fall in the cracks between teams.
- Exceptional written and verbal communication for delivering clear, data-driven, decision-oriented updates to executive stakeholders.
Preferred
- Experience defining or substantially redesigning a commissioning gate framework for site deployments.
- Experience coaching or developing more junior TPMs in technical depth and program execution.