Solutions Architect - CPU and LPU

NVIDIA
Shenzhen
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: mastersSkills: ["Communication","Problem-solving"]

Drive customer adoption of next-generation AI infrastructure spanning NVIDIA CPU platforms and LPU-based inference systems. Act as the first line of technical expertise to help design and optimize heterogeneous AI workloads across CPU, GPU, and LPU—improving latency, throughput, utilization, and cost. Build proof-of-concepts and reference architectures, tune LLM/generative AI pipelines across runtime and serving layers, and translate customer needs into platform feedback and roadmap input.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Solutions Architect - CPU and LPU

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Drive customer adoption of next-generation AI infrastructure spanning NVIDIA CPU platforms and LPU-based inference systems. Act as the first line of technical expertise to help design and optimize heterogeneous AI workloads across CPU, GPU, and LPU—improving latency, throughput, utilization, and cost. Build proof-of-concepts and reference architectures, tune LLM/generative AI pipelines across runtime and serving layers, and translate customer needs into platform feedback and roadmap input.
Location: Shenzhen
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Evangelize NVIDIA CPU platforms (Grace, Vera, and future generations) and LPU-based systems, focusing on AI software stacks and workload efficiency.
  • •Help customers design and optimize AI workloads across CPU, GPU, and LPU to improve latency, throughput, utilization, and cost efficiency.
  • •Analyze and tune LLM and generative AI pipelines across serving, runtime, memory, I/O, batching, scheduling, and orchestration layers.
  • •Build proof-of-concepts, reference architectures, and technical guidance with Engineering, Product, and Sales teams.
  • •Establish trusted technical relationships and provide strategic advisory for heterogeneous AI system design.

Key Requirements

  • •MS or PhD in Computer Science, Engineering, Mathematics, Physics, or a related field, or equivalent experience.
  • •5+ years in AI systems, infrastructure, performance engineering, or solution architecture.
  • •Strong understanding of modern CPU architecture, Linux systems, and software performance tuning.
  • •Hands-on experience in AI inference for LLM, generative AI, or agentic AI workloads.
  • •Experience optimizing heterogeneous systems (CPU and accelerators) with frameworks such as PyTorch, Triton, TensorRT-LLM, vLLM, or ONNX Runtime.
Experience:5+ yearsAI infrastructureAI inferenceHeterogeneous systemsLLM inferencePerformance engineering
Education:Master's
Skills:CommunicationProblem-solving
Tech Stack:LinuxPyTorchTritonTensorRT-LLMVLLMONNX RuntimeArm64

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor