Senior System Software Engineer, Agentic Inference - Dynamo

Nvidia
US, CA, Santa Clara / Remote - US2026-07-27remote_local

About the job

We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers for its GPU-accelerated deep learning software team. Academic and commercial groups around the world are using GPUs to power a revolution in AI, enabling breakthroughs in problems from image classification to speech recognition to natural language processing. We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier and accessible to all users.

Responsibilities

- In this role, you will develop open source software to serve inference of trained AI models running on GPUs.

- Contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads, including long-horizon reasoning, tool calling, and stateful, multi-turn execution.

- Innovate in inference-state management for long-running agents, including KV- and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL, to reduce repeated prompt processing, improve latency and token throughput, maximize GPU utilization, and lower per-token and per-task costs for self-hosted LLMs.

- Build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM, delivering day-0 support for new models, model-specific request parameters, upstream API compatibility, and stateful Responses API semantics.

- Balance a variety of objectives: build robust, scalable, high performance software components to support our distributed inference workloads; work with team leads to prioritize features and capabilities; load-balance asynchronous requests across available resources; optimize throughput under latency constraints; and integrate the latest open source technology.

Qualifications

Minimum

- Masters or PhD or equivalent experience

- 10+ years in Computer Science, Computer Engineering, or related field

- Ability to work in a fast-paced, agile team environment

- Excellent Rust/Python programming and software design skills, including debugging, performance analysis, and test design.

- Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs.

Preferred

- Prior contributions to open-source AI inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang).

- Experience optimizing GPU memory, KV and prefix caches, or high-performance networking for long-context, reasoning, and tool-calling workloads.

- Understanding of LLM-specific inference challenges for agentic workloads, including context and reasoning-token growth, bursty tool-call-driven traffic, multi-turn state reuse, and scheduling across concurrent trajectories.

- Prior experience integrating self-hosted LLM serving stacks with agent harnesses such as OpenCode, Codex, Claude Code, and Pi, including compatibility for APIs, streaming, structured outputs, tool calls, and session semantics.