Developer Technology Engineer - AI

NVIDIA
Shanghai, Shenzhen, Beijing
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 2+ yearsEducation: bachelorsSkills: ["Communication","Cooperation","Collaboration","Problem-solving","System architecture thinking"]

Build and optimize GPU-accelerated parallel algorithms for training and inference of large language models. Work with application developers and NVIDIA teams to improve architectures, software platforms, and programming models, turning real-world feedback into platform enhancements. Deeply optimize high-performance operators (GPU kernels, instruction tuning, compiler optimization) and advance distributed training/inference communication by improving communication libraries and data-transfer strategies for compute/communication overlap.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Developer Technology Engineer - AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build and optimize GPU-accelerated parallel algorithms for training and inference of large language models. Work with application developers and NVIDIA teams to improve architectures, software platforms, and programming models, turning real-world feedback into platform enhancements. Deeply optimize high-performance operators (GPU kernels, instruction tuning, compiler optimization) and advance distributed training/inference communication by improving communication libraries and data-transfer strategies for compute/communication overlap.
Location: Shanghai, Shenzhen, Beijing
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Partner with key application developers to understand current and future problems and deliver GPU-optimized solutions via library development and application contributions.
  • •Collaborate with architecture, research, libraries, tools, and system software teams to influence next-generation architectures, software platforms, and programming models.
  • •Optimize high-performance operators including GPU kernel optimization, instruction-level tuning, and compiler optimization across NVIDIA libraries and open-source projects.
  • •Advance distributed training and inference by refining communication libraries and open-source communication libraries.
  • •Study interconnect topologies and network protocols to design efficient data transfer strategies for compute-communication overlap.

Key Requirements

  • •A degree or equivalent experience in an engineering or computer science field; a masters or doctoral degree is preferred.
  • •2+ years of work experience.
  • •Solid understanding of C, C++, Python, or Fortran.
  • •Strong mathematical fundamentals including linear algebra and numerical methods.
  • •Background in parallel programming and accelerated computing, including performance analysis/tuning; GPU programming is desirable, plus distributed communication optimization experience is advantageous.
Experience:2+ yearsLarge language modelsHigh-performance computingDistributed trainingAccelerated computingOpen-source
Education:Bachelor's
Skills:CommunicationCooperationCollaborationProblem-solvingSystem architecture thinking
Tech Stack:CC++PythonFortranGPUsCUDAMegatronTRTLLMSGLangVLLMCuDNNCuBLASCUTLASSDeepGEMMFlashMLAFlashAttentionFlashinferNCCLGINNVSHMEM

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor