Deep Learning Solution Architect - Agentic Performance

NVIDIA
Beijing, Shanghai
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringEducation: mastersSkills: ["Communication","Stakeholder management"]

Architect production-grade generative AI solutions for enterprise customers, focusing on LLM pretraining, fine-tuning, high-performance inference, RAG workflows, and agentic inference orchestration using NVIDIA’s software and hardware ecosystem. Partner with customers to translate business challenges into tailored solutions, lead distributed optimization and performance tuning for throughput/latency/memory efficiency, and design RAG and agent pipelines for customer systems while supporting pre-sales technical activities like workshops and demos.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Deep Learning Solution Architect - Agentic Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Architect production-grade generative AI solutions for enterprise customers, focusing on LLM pretraining, fine-tuning, high-performance inference, RAG workflows, and agentic inference orchestration using NVIDIA’s software and hardware ecosystem. Partner with customers to translate business challenges into tailored solutions, lead distributed optimization and performance tuning for throughput/latency/memory efficiency, and design RAG and agent pipelines for customer systems while supporting pre-sales technical activities like workshops and demos.
Location: Beijing, Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Architect end-to-end solutions for LLM pretraining, fine-tuning, high-performance inference, RAG workflows, and agentic inference orchestration using NVIDIA platforms.
  • •Collaborate with customers to understand LLM-related business challenges and design tailored solutions aligned with the NVIDIA ecosystem.
  • •Lead LLM training, distributed optimization, and performance tuning to improve throughput, latency, and memory efficiency.
  • •Design and integrate RAG workflows and agentic inference pipelines into customer systems and provide best-practice guidance.
  • •Support pre-sales technical activities by collaborating with NVIDIA engineering teams on workshops and demos.

Key Requirements

  • •Master’s or Ph.D. in Computer Science or Artificial Intelligence, or equivalent experience.
  • •4+ years hands-on AI experience with open-source LLM training, fine-tuning, and production inference optimization.
  • •Proficiency with LLM architectures and LLM customization using PyTorch and Hugging Face Transformers.
  • •Strong knowledge of GPU computing, cluster architecture, and distributed parallel training/inference for LLMs.
  • •Competency in agentic inference design and using AI agents to solve business challenges.
Experience:Open-sourceLLM trainingDistributed trainingGenerative AI
Education:Master's in Computer Science / Artificial Intelligence
Skills:CommunicationStakeholder management
Tech Stack:PyTorchHugging Face TransformersRAGDockerKubernetesTRT-LLMMegatron-LMNVIDIA NeMoGPU computingDistributed parallel trainingQuantizationKV Cache tuningAgentic inference orchestration

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor