Senior Solutions Architect, Inference Service Providers

NVIDIA
Santa Clara, United States
Workplace: OnsiteFull timeUSD 184,000 - 356,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 6+ yearsEducation: bachelorsSkills: ["Technical leadership","Mentorship","Partner communication"]

Work with inference partners to help them deploy and operate distributed AI inference services at scale. Build and optimize inference “recipes” using NVIDIA Dynamo and accelerate pipelines with TensorRT-LLM, vLLM, and SGLang. Guide Kubernetes-based, disaggregated inference architectures through complex performance and reliability challenges, producing reference architectures and teaching partners and internal teams best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Solutions Architect, Inference Service Providers

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Work with inference partners to help them deploy and operate distributed AI inference services at scale. Build and optimize inference “recipes” using NVIDIA Dynamo and accelerate pipelines with TensorRT-LLM, vLLM, and SGLang. Guide Kubernetes-based, disaggregated inference architectures through complex performance and reliability challenges, producing reference architectures and teaching partners and internal teams best practices.
Location: Santa Clara, United States
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Work with inference partners to teach the value of the stack and bring back product insights for the industry.
  • •Build and operate inference recipes, distributing tasks among GPU workers to improve efficiency.
  • •Accelerate inference pipelines using TensorRT-LLM, vLLM, and SGLang to support disaggregated inference.
  • •Advocate and evangelize DevOps best practices for managing Kubernetes clusters, compute fabrics, and observability.
  • •Provide mentorship and technical leadership to customers and internal teams to deploy disaggregated inference systems and resolve complex issues.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •6+ years in solutions architecture (or similar) driving customer engagements to deploy distributed systems, including 2+ years with AI workloads on Kubernetes.
  • •Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization and serving.
  • •Deep knowledge of inference best practices, including disaggregated serving, KV cache management, speculative decoding, quantization, and custom inference kernels.
  • •Hands-on full-stack agent design covering modern sandboxing, memory and retrieval systems with evaluation, skill design, and governance.
  • •BS in CS/Engineering or equivalent experience.
Experience:6+ yearsAIKubernetes
Education:Bachelor's in CS/Engineering
Skills:Technical leadershipMentorshipPartner communication
Tech Stack:NVIDIA DynamoTriton Inference ServerTensorRT-LLMVLLMSGLangKubernetesObservabilityKV cache managementSpeculative decodingQuantizationCustom inference kernelsSFTDPOGRPORLVRNIXLGroveNVIDIA Cloud Partner

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor