Senior Research Engineer - Enterprise Products

Nvidia
US, WA, Remote / US, CA, Santa Clara / Remote - US2026-07-10remote_local

About the job

We are now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited to change the way people infuse AI into products and services? NVIDIA is at the forefront of generative AI models, from language to images. NVIDIA provides building blocks to democratize AI and make generative AI easy to develop, integrate, and deploy. Our team is dedicated to developing optimized inferencing technologies to support our growing generative AI needs. We contribute to all steps of the machine learning lifecycle: from conceptualization, to applied research, engineering for optimized inference, and deployment. Collaborate with research teams, engineers, and open-source community.

Responsibilities

Design and evaluate routing policies for LLM traffic to best use mixture of model systems.

Build and run agentic benchmarks (e.g., Terminal-Bench ) to measure algorithm quality, and turn results into calibration data and routing profiles

Ship to an open-source repo: design docs, code review, docs, and community contributions

Collaborating with engineering teams across all of NVIDIA to ensure our software integrates seamlessly up and down the NVIDIA accelerated serving stack.

Qualifications

Minimum

Bachelor's of Master's degree in Computer Science or equivalent experience.

8+ years of industry experience in Deep Learning frameworks (PyTorch or TensorFlow).

Experience designing or running LLM evaluations/benchmarks — ideally agentic ones — and drawing statistically sound conclusions from them

Understanding of modern techniques in Machine Learning, Deep Neural Networks, Natural Language Processing, or Speech Recognition.

Empirical research mindset: forming hypotheses about new algorithms, running calibrations, iterating on results

Strong communication and interpersonal skills, along with the ability to work in a dynamic and distributed team. A history of mentoring junior engineers and interns is a huge plus.

A desire to constantly grow and learn new things.

Strong computer science fundamentals - algorithms and data structures, computational complexity, parallel and distributed computing, system software.

Preferred

Experience architecting or developing large-scale distributed systems for deep learning.

Agentic benchmark creation and publications.

Knowledge of CPU and/or GPU architecture.

GPU programming (CUDA).