About the job
OCI (Oracle Cloud Infrastructure) AI Infrastructure is at the forefront of building a cutting-edge, ultra-high-performance GPU platform designed to support AI/ML/HPC workloads. This is your chance to be part of the AI revolution, creating systems that allow customers to scale from tens to thousands of GPUs without compromising performance.
Our team is responsible for designing and developing fundamental architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services. These are essential for running distributed AI/ML/HPC workloads across thousands of GPUs, leveraging technologies like RoCE and InfiniBand.
We're looking for an experienced front-line engineering manager to lead and support AI Data Plane (DP) team. You'll build a highly available, massive scale, integrated cloud service in a distributed, multi-tenant cloud environment for hyper-scale AI customers.
You'll bring experience in leading engineering teams, including hands-on technical management and on-call experience. You have technical experience with high-performance, mission-critical environments. You demonstrate strong ownership, solid communication skills and a bias for action. You're comfortable leading projects, communicating with senior leadership, and growing your team.
Responsibilities
Engineering and project management: Work with customers and stakeholders to understand feature requirements and pain points, build roadmaps, make priority tradeoffs, plan and drive engineering projects, manage risks and blockers, and break large technically complex projects into manageable pieces.
People management: Handle coaching and performance-management for engineers, forecast staffing needs, and work with recruiting teams to hire positions.
Process improvement: Work across engineering teams to identify and resolve systemic problems and bottlenecks to improve overall engineering velocity and efficiency.
Lead and support the AI Data Plane (DP) team, building a highly available, massive scale, integrated cloud service in a distributed, multi-tenant cloud environment.
Collaborate cross-functionally with PM team and other OCI Core Service teams (Compute, Storage, Network).
Qualifications
Minimum
4+ years of engineering management and people management experience
5+ years in leading large cross-functional projects, operating large-distributed services in cloud environments
5+ years of owning/driving roadmap strategy and definition
Bachelor's degree in computer science or related engineering field
Strong technical knowledge in distributed systems, high performance computing, and GPU systems
Proven industry expert in high scalability, availability, low-latency domains
Excellent organizational, verbal, and written communication skills
Experience working in cloud platform(s) (AWS, OCI, GCP, Azure etc)
Preferred
No preferred qualifications listed.