About the job
You’ll be working in the compute team focusing on GPU workload scheduling and inference serving optimization. You would partner with the inference team to improve our inference throughput and latency for evals and reinforcement learning. You would collaborate with our scalability team to focus on stabilizing our large scale fault tolerant training. You would also be in close contact with the infra team to make sure our GPU nodes are all healthy and fully utilized.
Responsibilities
- Design and develop internal scheduling system to maximize GPU utilization
- Build API and tooling to help manage the lifecycle of GPU workloads and troubleshoot failures
- Design and improve inference control plane to speed up model deployment and inference request serving
- Collaborate with research to improve research velocity continuously
Qualifications
Minimum
- Strong programming skills in Go, or other similar languages
- Strong systems engineering background: distributed systems, schedulers, control planes, or high-throughput data planes.
- Production experience with Kubernetes internals — controllers, informers, operators — not just deploying to it.
- Bias toward observability and debuggability: building a system that is easy to navigate when debugging production issues
Preferred
- Plus: experience in systems serving large scale inference requests