About the job
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing open problems in autonomous driving, towards the goal of safely operating Waymo vehicles in dozens of cities and under all driving conditions. In this hybrid role, you will report to a Senior Director of AI Foundations
Responsibilities
Own the data recipe for Waymo’s Foundation Model pre-training and post-training
Lead and drive science on best practices around clustering, filtering, de-duplication, long-tail data mining, memorization, etc.
Tech-lead a team of research engineers to build the data flywheel to enable the above from Waymo’s massive driving data
Integrate emerging research from the broader community to do rigorous ablations and promising data research techniques. This includes scaling ladders for pre-training, data mix optimizations and RL recipes (preferences) for post-training.
Engage with the wider research community on best practices around data centric evaluation creation and refinement
Partner with engineering and research teams across Waymo to share recipes, techniques, and post-training best practices to accelerate our collective know-how
Qualifications
Minimum
PhD or Masters in Computer Science, Machine Learning, Robotics, or a similar technical field; with 3+ years of industry or post-doc research experience in Data-centric AI, Reinforcement Learning or Foundation Models
Demonstration of original contributions to the field through high-impact publications (ArXiv, peer-reviewed conferences like NeurIPS/ICLR/CVPR), technical blog posts, or significant open-source contributions
Proficiency and in-depth knowledge of the inner workings of an ML framework (e.g. Pytorch, JAX, Tensorflow)
Preferred
Extensive experience working with data and recipes for large scale foundation models
ML infra experience: training, evaluating and deploying ML models at scale
Deep learning experience, especially with generative models, e.g., LLMs/VLMs, and/or reinforcement learning