About the job
Wayve is building autonomous driving technology that runs on real vehicles. Getting our models onto embedded hardware — correctly, quickly, and reproducibly — is one of the hardest problems between research and product.
As a ML Compiler Engineer, you will own the compilation pipeline that makes that possible. You will build and extend Wayve's ML compiler end-to-end: designing passes, integrating with vendor toolchains like NVIDIA TensorRT and Qualcomm QNN, and delivering deployable bundles that meet our accuracy and latency requirements on every target platform.
Each stage in the pipeline — capture, decomposition, precision assignment, legalisation, partitioning — can affect accuracy, latency, or whether a vendor backend accepts the graph. Your work spans the full lowering stack, building compiler passes and infrastructure that scale across architectures and target platforms.
Responsibilities
Own the ML compilation pipeline end-to-end — from checkpoint to deployable bundle on NVIDIA (TensorRT) and Qualcomm (QNN) targets.
Design and implement compiler passes with accuracy and latency gates, so bad compiles are caught before they reach hardware.
Build compilation infrastructure that scales across platforms, model architectures, and SoCs — without re-engineering for each new target.
Partner with model and training teams on compilability; build regression and benchmarking to validate changes across releases.
Set technical direction and raise the bar for compiler engineering across the team.
Qualifications
Minimum
Built or owned significant parts of an ML compilation or graph-lowering pipeline.
Deep experience with quantisation in compilation — precision typing, PTQ integration, debugging accuracy loss from compiler transforms.
Strong Python; comfortable building and testing compiler infrastructure in production codebases.
Proficiency with at least one of: MLIR, ONNX, TensorRT, Qualcomm QNN, PyTorch graph capture/export.
Experience with multi-target compilation or graph partitioning across hardware backends.
Ability to reason about correctness and performance trade-offs at each compiler stage.
Preferred
C++ a plus.