Senior Product Manager – AI Inference Performance

NVIDIA
Santa Clara
Full timeUSD 208,000 - 327,750 annuallyFunction: Product ManagementExperience: 12+ yearsSkills: ["Independent execution","Strategic thinking","Data-driven decision-making","Translating technical capability into business value","Operational rigor"]

Own NVIDIA’s inference performance product roadmap, setting direction across the inference stack from model representation and memory/state management to request scheduling and token generation. Build scalable capabilities that generalize across model families and deployment topologies. Define strategies for agentic multi-turn workloads and framework/ecosystem rollouts across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo, while establishing benchmark methodology and running daily release readiness and feedback loops.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 days ago

Senior Product Manager – AI Inference Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 28 minutes agoStatus: Live

Job Summary

Own NVIDIA’s inference performance product roadmap, setting direction across the inference stack from model representation and memory/state management to request scheduling and token generation. Build scalable capabilities that generalize across model families and deployment topologies. Define strategies for agentic multi-turn workloads and framework/ecosystem rollouts across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo, while establishing benchmark methodology and running daily release readiness and feedback loops.
Location: Santa Clara
Employment Type: Full time
Job Function: Product Management

Key Responsibilities

  • •Own the inference performance roadmap and set direction across how models are represented, memory/state managed, requests scheduled and served, and tokens generated.
  • •Build broadly reusable platforms and capabilities that generalize across model families, deployment topologies, and customer sizes.
  • •Define performance strategy for agentic and multi-turn workloads, including cross-turn cache reuse, request prioritization, and idle-time handling.
  • •Set framework and ecosystem strategy for how optimizations land across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.
  • •Own benchmarking and performance claims, including methodology, key metrics (TTFT, ITL, throughput per GPU, cost per million tokens), and credibility guardrails.

Pay and Benefits

Salary: USD 208,000 - 327,750 annually
Equity and Bonus:Equity

Key Requirements

  • •12+ years in product management at a technology company, or comparable time as a founder, engineering lead, or technical product owner.
  • •Deep knowledge of AI inference optimization, including KV caching/reuse, quantization, speculative decoding, and disaggregated serving.
  • •Experience with inference and orchestration frameworks such as TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, and related serving/scheduling ecosystems.
  • •Proven ability to work independently to define strategy for ambiguous problems and drive shipped outcomes.
  • •BS, MS, or PhD in Computer Science, Computer Engineering, or a relevant field (or equivalent experience).
Experience:12+ yearsAI inferenceDeep learningGenerative AILLM inference
Education:
Skills:Independent executionStrategic thinkingData-driven decision-makingTranslating technical capability into business valueOperational rigor
Tech Stack:TensorRT-LLMVLLMSGLangNVIDIA DynamoKV cachingQuantizationSpeculative decodingDisaggregated servingTriton Inference ServerTTFTITLBenchmark methodologyProfilingKernel-level optimizationInference and orchestration frameworksRelease managementRegression trackingAutoscalingSLA management

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor