ML Researcher - Image / Video Diffusion

Krea
San Francisco, CA, USA2026-09-01OnSite

About the job

We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale.

Responsibilities

- Train diffusion models for image and video generation on large GPU clusters.

- Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.

- Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.

- Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.

- Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.

- Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.

Qualifications

Minimum

- Proven track record in working with image or video models at scale.

- Strong proficiency in PyTorch and understanding of its inner workings.

- Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP.

- Experience in profiling and debugging large distributed training.

- Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.

- Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.

- Being comfortable working in a goal-oriented research environment.

- Comfortable working with underspecified goals.

Preferred

- Publications or open-source contributions in image or video models at scale.

- Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.

- Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.

- Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.

- Ability to iterate rapidly, and propose creative research directions.

- Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.