Engineering Manager - Machine Learning

Recursion
Toronto, Canada2026-07-16

About the job

You will lead a team working to build, scale, and optimize the machine learning infrastructure that powers Recursion's drug discovery platform. From model training pipelines to production deployment systems, to agent infrastructure and Large Language Models, you will ensure our ML models can operate at massive scale across our supercomputing infrastructure, both on prem and in the cloud. You will work cross-functionally across ML engineering, data science, and research teams to translate requirements into robust, scalable ML infrastructure solutions.

Responsibilities

Enable AI/ML, LLM, and Agentic Systems teams for scale - The ML infrastructure team is responsible for building and operating platforms that allow data scientists and ML engineers to train, deploy, and monitor models across Recursion's massive datasets.

Act as a mentor, coach, and sponsor - You will share your technical, leadership and managerial skills in MLOps, distributed computing, and infrastructure engineering, delivering impact, learning, and growth across teams at Recursion.

Enable a model-driven culture - Machine learning is at the core of everything we do. You will work with stakeholders across the business to ensure our ML infrastructure supports rapid experimentation, reliable model deployment, and continuous improvement.

Qualifications

Minimum

Experience in a hands-on technical role as a tech lead or a manager with a focus on infrastructure, MLOps and distributed systems.

A people-first mindset.

Demonstrated past record of learning from and teaching peers in areas of ML infrastructure, model deployment, distributed compute, GPU optimization, and MLOps system architecture

Excitement to learn parts of our ML tech stack that you might not already know.

Preferred

Fluency in life sciences or drug discovery is a plus but not required to be considered.