About the job
Inference is the fastest growing and most competitive area in Generative AI today. It is where AI models impact our daily life, and where ever bit of accuracy and performance matters for quality, safety, and cost. Inference is also constantly evolving, with new acceleration algorithms, usecases, and deployment techniques. As a Senior Product Manager for AI Platform Inference you will be responsible for building the tools, SDKs, and libraries which enables developers' Inference deployments to thrive on NVIDIA GPUs.
Responsibilities
Create products to help developers build better Inference deployments
Develop product strategy, roadmaps, and go-to-market plans
Collaborate with internal and external developers to build product-based roadmaps for model optimization software
Work with leadership to align with and drive company strategy
Qualifications
Minimum
Experience with Inference deployment and optimization software (ex. vLLM, SGLang, FlashInfer, TensorRT-LLM, Triton, Dynamo, TorchAO, etc.)
Demonstrable knowledge of GenAI or machine learning concepts, particularly around performance optimization, and software development and delivery
BS or MS degree in Computer Science, Computer Engineering, or similar experience (or equivalent experience)
6+ years of technical product management, or similar, experience at a technology company
Strong communication and interpersonal skills
Preferred
Experience leading optimization products for Inference
Working on Open Source & Github-first developer products with deep customer interactions
Knowledge of GPU architecture, HW/SW co-design, and performance profiling