About the job
We’re hiring a Platform Software Engineer to help build the execution layer behind OCI’s AI and GPU growth motion. You’ll work across Oracle data platforms, GPU scheduling, orchestration pipelines, and MLOps agentic tooling — turning customer demand into production-grade systems that scale across our cluster fleet. This is a builder role on a small, high-leverage team.
Responsibilities
• Integrate with Oracle data platforms (OCI Streaming, 23ai, Object Storage, Data Flow) to move training and inference data reliably across customer and internal pipelines.
• Build and extend GPU schedulers and capacity-aware placement logic for A100, H100, H200, and Blackwell fleets.
• Develop orchestration pipelines for training, fine-tuning, and inference workloads using Kubernetes, Argo, and Slurm where appropriate.
• Ship MLOps agentic tooling — observability, automated triage, cost and SLO agents — that reduces operator load on large GPU deployments.
• Partner with PMs, SAs, and customer-facing teams to convert field requirements into reusable components, not one-off scripts
Qualifications
Minimum
• 3–6 years of software engineering experience, ideally in distributed systems, data platforms, or ML infrastructure.
• Strong Python; working knowledge of Go or Rust is a plus.
• Hands-on experience with Kubernetes, container runtimes, and at least one workflow engine (Argo, Airflow, Prefect, Temporal).
• Familiarity with GPU workloads, CUDA toolchains, or inference serving frameworks (vLLM, TensorRT-LLM, Triton).
• Experience with LLM agent frameworks (LangGraph, CrewAI, or similar) is a strong plus.
• Comfort working in ambiguity, shipping iteratively, and reasoning about topology — not just features.
Preferred
• Prior work on multi-cloud or hybrid deployments.
• Exposure to QEC, quantum-classical hybrid workloads, or scientific computing.
• Open-source contributions to the AI infra ecosystem.