Inference Performance Engineer, AI Inference Configuration Optimization

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 124,000 - 195,500 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Experimental methodology","Communication","Optimization","Benchmarking","Collaboration"]

Drive performance improvements for large-scale AI inference by designing and validating autonomous, evidence-backed optimization workflows. Explore configuration options to increase throughput-per-GPU and interactivity while handling batching, KV cache, quantization, and speculative decoding. Benchmark and profile across serving architectures and platforms using Nsight tools, CUPTI, and profiling analysis, then upstream serving patches, optimized kernels, and deployment recipes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
16 hours ago

Inference Performance Engineer, AI Inference Configuration Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Drive performance improvements for large-scale AI inference by designing and validating autonomous, evidence-backed optimization workflows. Explore configuration options to increase throughput-per-GPU and interactivity while handling batching, KV cache, quantization, and speculative decoding. Benchmark and profile across serving architectures and platforms using Nsight tools, CUPTI, and profiling analysis, then upstream serving patches, optimized kernels, and deployment recipes.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Build reusable, agent-executable optimization workflows by distilling performance methods, reviewing agent experiments, and curating best-known configurations.
  • •Improve AI inference workload performance by exploring configuration options such as parallelism, batching, KV cache handling, quantization, and speculative decoding settings.
  • •Measure and optimize aggregated and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA GPU platforms.
  • •Profile workloads using Nsight Systems, kernel traces, and internal analysis tools, using roofline and speed-of-light analysis to validate hypotheses and deliver measured wins.
  • •Collaborate with serving framework, kernel, benchmarking, and GPU architecture teams to upstream patches, optimized kernels, and deployment recipes while maintaining model correctness.

Pay and Benefits

Salary: USD 124,000 - 195,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience.
  • •3+ years of relevant engineering experience.
  • •Extensive knowledge of AI model execution efficiency, including continuous batching, throughput-latency tradeoffs, KV cache/memory limits, parallelism, MoE serving, quantization, and meeting serving SLAs.
  • •Hands-on benchmarking and profiling of GPU workloads using tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler, and interpreting kernel-level performance data.
  • •Strong Python engineering skills and ability to navigate and modify large C++/CUDA serving codebases, using rigorous, reproducible experimental methodology.
Experience:3+ yearsAI inferenceLarge-scale AIGPU workloadsModel servingAgentic AI
Education:
Skills:Experimental methodologyCommunicationOptimizationBenchmarkingCollaboration
Tech Stack:PythonC++CUDATensorRT-LLMSGLangVLLMDynamoFlashInferNsight SystemsNsight ComputeCUPTIPyTorch profilerTensor CoresTMANCCLNIXLNVSHMEMMLPerf InferenceSemiAnalysis InferenceXRoofline

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor