Senior Deep Learning Solution Architect

NVIDIA
Beijing, Shanghai
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringEducation: mastersSkills: ["Learning agility","Adaptability","Independent problem-solving","Performance analysis","Technical exploration"]

Contribute to open-source inference frameworks such as SGLang and vLLM, building features and operators while optimizing performance and adding model support with the community. Develop KV cache offloading frameworks for LLM workloads across CPU, SSD, and remote storage to improve inference efficiency. Drive R&D on compute performance for distributed training, analyze ML system bottlenecks, and create acceleration libraries or frameworks accordingly.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 weeks ago

Senior Deep Learning Solution Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Contribute to open-source inference frameworks such as SGLang and vLLM, building features and operators while optimizing performance and adding model support with the community. Develop KV cache offloading frameworks for LLM workloads across CPU, SSD, and remote storage to improve inference efficiency. Drive R&D on compute performance for distributed training, analyze ML system bottlenecks, and create acceleration libraries or frameworks accordingly.
Location: Beijing, Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Contribute to open-source inference frameworks (SGLang, vLLM) including feature/operator development, performance optimization, and model support.
  • •Develop and optimize KV cache offloading frameworks for LLM workloads, enabling multi-level cache offloading and reuse across CPU, SSD, and remote storage.
  • •Drive R&D on compute performance for distributed training and explore performance optimization methods and technologies.
  • •Analyze computational challenges in ML systems to identify bottlenecks and build example code, acceleration libraries, or frameworks.
  • •Define new technical problems and explore solutions, including research into performance and system acceleration.

Key Requirements

  • •Over 5 years of technology industry experience with a master’s degree or above in computer science, mathematics, electrical engineering, automation, or related fields.
  • •Strong interest in accelerated, parallel, and heterogeneous computing, with motivation to explore these areas deeply.
  • •Solid programming skills with a good understanding of data structures and computer systems fundamentals.
  • •Strong learning agility and the ability to independently analyze, define, and explore technical problems.
  • •Ability to build acceleration libraries/frameworks and strong learning agility; proficiency with AI coding tools.
Experience:AI computingHPCLLM inferenceDistributed trainingOpen-source software
Education:Master's in computer science, mathematics, electrical engineering, automation, or related fields
Skills:Learning agilityAdaptabilityIndependent problem-solvingPerformance analysisTechnical exploration
Tech Stack:SGLangVLLMKV cache offloadingFlexKVCPUSSDRemote storageDistributed training

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor