Senior Deep Learning Systems Architect

Nvidia
US, CA, Santa Clara2026-07-15onsite

About the job

NVIDIA is seeking architects like you to help design hardware accelerator and processor architectures that enable state of the art machine learning and data analytics algorithms and applications on our next-generation mobile, embedded and datacenter platforms. This position offers you the opportunity to have a real impact in a dynamic, technology-focused company.

Responsibilities

Contribute to features that help next-generation GPUs and systems advancing the state of AI.

Keep up with the latest DL research and collaborate with diverse teams (internal and external to NVIDIA), including DL researchers, hardware architects, and software engineers.

Participate in engineering projects and co-design architecture for systems from conception, specification and prototyping.

Understand various AI/DL workloads and their mapping to underlying HW and Systems. Identify potential improvements and bottlenecks, propose solutions to address existing gaps, and accelerate/improve current systems/methods.

Perform comprehensive analyses from first principles of various deep learning techniques, system optimizations to build out analytical models as well as implementing prototypes, and benchmarking to test/prove ideas.

Qualifications

Minimum

MS (or equivalent experience) or PhD degree in computer science, computer architecture, electrical engineering or related field with 10+ years of relevant work experience.

Strong background in at least a few of the following relevant areas: Machine learning (with focus on Deep Neural Networks), including a solid understanding of DL fundamentals; Experience adapting and training DNNs for various tasks; Experience developing code for one or more of the DNN training frameworks (such as PyTorch, TensorFlow or JAX): Numerical analysis, Performance analysis and optimization & Computer architecture.

Programming fluency with C++ and ideally Python.

Preferred

Work experience with GPU computing (CUDA, OpenCL, OpenACC) and HPC (MPI, OpenMP) is a huge plus.