Principal Engineer, CAPE

Crusoe
San Francisco, CA - US2026-07-21OnSite

About the job

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. This is a Principal Engineer role reporting into Cloud Availability, working directly on one of the highest-priority technical charters in the org.

Responsibilities

Build a unified observability plane that correlates GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals.

Treat tens of thousands of accelerators across sites as a single programmable system with one health model, one scheduler, and one source of truth.

Implement closed-loop autonomy to diagnose, decide, and remediate with no human in the loop — drain, checkpoint, replace, and resume a live job automatically.

Continuously maximize goodput by trading scheduling, placement, and maintenance decisions against it as the north-star metric.

Predict failures hours ahead by forecasting GPU, NVLink, optics, and thermal degradation before it stalls a job, and pre-emptively migrate work.

Detect stragglers and silent failures by isolating the exact rank, GPU, and node from collective-operation signals and surfacing root cause fast.

Schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom.

Build a digital twin of the fleet to simulate failures, scheduling policies, and remediation logic before they touch production.

Develop agentic operations that propose and execute fixes with guardrails and a full audit trail.

Implement zero-trust, fully auditable multi-tenancy where every action is identity-scoped, policy-checked, and streamed to customer audit systems in real time.

Create self-qualifying hardware pipelines where new and repaired nodes prove themselves through automated burn-in before taking customer load.

Qualifications

Minimum

10+ years building infrastructure-layer systems at scale — fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation.

Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and systems that make autonomous decisions against live production infrastructure.

Hands-on fluency with GPU/HPC infrastructure — GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior at the hardware level.

Track record of designing and shipping large-scale observability or telemetry platforms that correlate signals across compute, network, and storage layers.

Comfort operating in ambiguity and defining the architecture and standards for a system that doesn't exist yet — this is a 0→1 charter, not a maintenance role.

Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) and the judgment to know when to build vs. adopt existing tooling.

Preferred

Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus.

Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus but not required.