ML Infrastructure Engineer

xAI
Palo Alto, California, United States / Palo Alto, CA, Palo Alto, California, United States2026-07-21

About the job

As an ML Infrastructure Engineer, you will play a pivotal role in building and optimizing the reliable, high-performance ML platform that powers recommendations on X. We're looking for exceptional engineers who are passionate about our mission and have a strong desire to make a meaningful impact.

Responsibilities

Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses

Developing data pipelines and integrating large-scale data, training, and inference systems

Collaborating with ML teams to productionize models and ensure seamless integration across the stack

Ensuring scalability, reliability, and efficiency of large-scale machine learning systems

Working across the full stack to solve complex problems independently

Mentoring junior engineers and contributing to the growth of the team

Qualifications

Minimum

Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience

2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications

2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists

Strong proficiency with Python and experience with compiled languages such as C++ or Rust

Preferred

Deep familiarity with modern ML frameworks such as JAX or PyTorch

Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking

Comfortable with Linux systems and orchestration tools

Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling