Senior Machine Learning Engineer, Agent Oversight

Scale AI
San Francisco / New York / Seattle2026-07-14

About the job

As a Machine Learning Engineer on Agent Oversight, you will drive the end-to-end lifecycle that ensures our production agents perform reliably and improve over time. This includes building observability tools, designing robust evaluation frameworks, and developing improvement loops. Whether scaling infrastructure or researching new improvement methods, you will navigate the entire ML loop while maintaining rigorous technical standards.

Responsibilities

Build or contribute to observability into agent behavior in production — the signals and instrumentation needed to actually see what an agent is doing, not just whether it succeeded or failed

Design evaluation methodologies and metrics for agentic applications, and work with the platform to make them run automatically, at scale, across different customer use cases, not just as one-off analyses

Build, ship, and own ML systems that detect drift, anomalies, or misalignment in production agent behavior — from first prototype through running reliably at scale

Design and run rigorous experiments to validate model and agent performance improvements before they ship

Work alongside software engineers on the platform where your work intersects with broader infrastructure — but you’re expected to take your own work from idea to production, not hand it off

Collaborate closely with product managers, customers, data annotators, Forward Deployed Engineers, and other engineering teams to translate enterprise and government requirements into robust platform capabilities

Depending on focus, contribute to novel methods and approaches that push the state of the art for agent evaluation and improvement, or focus on building ML systems that hold up reliably at scale in production

Qualifications

Minimum

5+ years of experience as an ML engineer or applied scientist, ideally on a production ML or LLM-powered system — not just consuming a third-party ML API within a feature

Strong grounding in at least two of the following: Building or scaling evaluation, monitoring, or continuous-learning infrastructure for ML/agentic systems; Design experience for agent systems (architecture, orchestration, tool use); Developing new methods, reward models, or model training/fine-tuning approaches

Hands-on experience with LLMs and agent architectures — tool use, planning, multi-agent orchestration

Comfortable partnering with software engineers to productionize research and experimental work, not just deliver a one-off analysis

Rigorous approach to experimentation: clear hypotheses, real statistical grounding, and results that hold up under scrutiny

Track record of collaborating across functions (Product, Forward Deployed Engineering, etc.) to navigate ambiguous requirements and bring them to production

Gives direct, substantive feedback on designs and code, and takes it the same way — and mentors others as they grow

Preferred

Experience building or contributing to RLHF, SFT, or other fine-tuning/RL workflows, reward modeling, or verifiable-reward systems

Experience with model or systems optimization (e.g., latency, cost, or inference efficiency)

Published research, open-source contributions, or patents in agentic systems, LLMs, or applied ML

Experience working in regulated or enterprise contexts

Track record of taking a novel method from prototype to something running reliably in production, navigating ambiguity along the way

Experience reviewing others’ technical designs or mentoring engineers at a senior/staff level