Senior Software Engineer, AI/ML Platform

Agility Robotics
Salem / Pittsburgh2026-08-20

About the job

Join the team building the machine learning platform to power fleet-scale humanoid robotics. As a senior engineer on the ML Infrastructure and Platform group, you will help architect and build the foundational infrastructure for AI and machine learning operations at Agility. This includes the platform layer for data collection and processing, training, sim and real evaluation, and model management and observability.

Responsibilities

Contribute to the design and implementation of the ML platform for orchestrating the end to end AI flywheel of data processing, training, evaluation, and deployment

Develop reliable workflows across cloud compute, Kubernetes, and continuous automation

Build core infrastructure components such as the model registry, feature store and experiment tracking tooling

Own developer-facing APIs and CLI tools that make ML workflows simple and reproducible

Implement the CI/CD lifecycle for ML that enable continuous retraining, automated testing, and seamless model delivery to production environments

Work closely with the Staff ML Infra Engineer and cross-functional stakeholders (AI researchers and robotics engineers) to understand requirements and translate them into scalable solutions/systems

Partner with data platform engineers to integrate ML orchestration and metadata tracking tools with our existing data lake and pipelines

Apply MLOps best practices: reproducibility, lineage, rollback, monitoring and governance

Mentor junior engineers and influence the broader cloud platform organization’s roadmap

Qualifications

Minimum

5+ years of software engineering experience, with at least 2+ years working on ML infrastructure, data platforms or MLOps systems in production environments

Experience building and maintaining components of modern ML platforms—such as experiment tracking, model registries, training pipelines, or deployment systems

Familiarity with orchestration and tracking tools (MLflow, WandB, Airflow, Kubeflow, etc.)

Proficiency with cloud-native platforms (AWS, GCP, or Azure), containers, and IaC (e.g., CDK, Terraform)

Hands-on experience with processing or modeling multimodal data(sensor logs, camera streams, behaviour traces etc)

Comfortable collaborating cross-functionally with research scientists, data engineers, and robotics/autonomy teams to ship infrastructure used by others

Preferred

Experience with robotics, autonomous vehicles, drones or embedded ML

Contributions to open-source ML infrastructure or MLOps tooling a plus