Senior Core Infrastructure Engineer

Oracle
US2026-07-30

About the job

We’re hiring a Platform Software Engineer to help build the execution layer behind OCI’s AI and GPU growth motion. You’ll work across Oracle data platforms, GPU scheduling, orchestration pipelines, and MLOps agentic tooling — turning customer demand into production-grade systems that scale across our cluster fleet. This is a builder role on a small, high-leverage team.

Responsibilities

• Integrate with Oracle data platforms (OCI Streaming, 23ai, Object Storage, Data Flow) to move training and inference data reliably across customer and internal pipelines.

• Build and extend GPU schedulers and capacity-aware placement logic for A100, H100, H200, and Blackwell fleets.

• Develop orchestration pipelines for training, fine-tuning, and inference workloads using Kubernetes, Argo, and Slurm where appropriate.

• Ship MLOps agentic tooling — observability, automated triage, cost and SLO agents — that reduces operator load on large GPU deployments.

• Partner with PMs, SAs, and customer-facing teams to convert field requirements into reusable components, not one-off scripts

Qualifications

Minimum

• 3–6 years of software engineering experience, ideally in distributed systems, data platforms, or ML infrastructure.

• Strong Python; working knowledge of Go or Rust is a plus.

• Hands-on experience with Kubernetes, container runtimes, and at least one workflow engine (Argo, Airflow, Prefect, Temporal).

• Familiarity with GPU workloads, CUDA toolchains, or inference serving frameworks (vLLM, TensorRT-LLM, Triton).

• Experience with LLM agent frameworks (LangGraph, CrewAI, or similar) is a strong plus.

• Comfort working in ambiguity, shipping iteratively, and reasoning about topology — not just features.

Preferred

• Prior work on multi-cloud or hybrid deployments.

• Exposure to QEC, quantum-classical hybrid workloads, or scientific computing.

• Open-source contributions to the AI infra ecosystem.