Principal Architect, AI Inference Networking - NIXL and Dynamo

NVIDIA
Shanghai, Shenzhen
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 12+ yearsSkills: ["Communication","Technical credibility","Community building","Cross-time-zone collaboration"]

Architect and drive adoption of NIXL in China, partnering with platform teams to diagnose inference bottlenecks at scale and translate findings into upstream code and roadmap. Build NIXL backends and plugins, extend Dynamo integrations, and help define dynamic data-path APIs for elastic, low-latency serving. Profile KV-cache and tensor/GPU transfer performance, contribute optimizations, and grow the local NIXL community through talks and reference architectures.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 days ago

Principal Architect, AI Inference Networking - NIXL and Dynamo

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Architect and drive adoption of NIXL in China, partnering with platform teams to diagnose inference bottlenecks at scale and translate findings into upstream code and roadmap. Build NIXL backends and plugins, extend Dynamo integrations, and help define dynamic data-path APIs for elastic, low-latency serving. Profile KV-cache and tensor/GPU transfer performance, contribute optimizations, and grow the local NIXL community through talks and reference architectures.
Location: Shanghai, Shenzhen
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Drive NIXL adoption in China from architecture conversations through prototype, benchmarks, production rollout, and upstream contributions.
  • •Profile customer inference stacks and fix throughput and TTFT issues (KV cache movement, memory registration, transfer scheduling).
  • •Write NIXL backends and plugins, extend Dynamo integrations, and contribute optimizations upstream to NIXL and GPUDirect-class data paths.
  • •Help define dynamic data-path APIs supporting elastic scaling, worker churn, and runtime rerouting in production inference clusters.
  • •Bring regional requirements into NVIDIA’s roadmap and grow the local NIXL community through talks, benchmarks, and reference architectures.

Key Requirements

  • •M.Sc. or Ph.D. in CS/EE/CE, or equivalent depth earned in industry.
  • •12+ years building or optimizing large-scale distributed systems for AI inference/training infrastructure, communication libraries, high-performance networking, or HPC runtimes.
  • •Strong systems programming in C++ and Python, with performance-critical code shipped.
  • •Hands-on experience with RDMA networking (InfiniBand or RoCE) and GPU memory movement.
  • •Working knowledge of modern LLM serving, including prefill/decode disaggregation, KV cache management, tensor/pipeline parallelism, and at least one serving stack (Dynamo, TensorRT-LLM, vLLM, SGLang, Triton).
Experience:12+ yearsAI inferenceAI trainingDistributed systemsHPCLLM servingOpen sourceNetworkingGPU
Education:
Skills:CommunicationTechnical credibilityCommunity buildingCross-time-zone collaboration
Languages:MandarinEnglish
Tech Stack:C++PythonNIXLNVIDIA DynamoGPUDirectCUDARDMAInfiniBandRoCELLM servingKV cacheTensor parallelismPipeline parallelismDynamoTensorRT-LLMVLLMSGLangTritonNCCLMooncake

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor