Senior Deep Learning Software Engineer, Inference

NVIDIA
United States
Workplace: RemoteFull timeUSD 152,000 - 287,500Function: Software EngineeringExperience: 5+ yearsSkills: ["Performance optimization","Analysis","Tuning","Software design","Cross-collaboration","Debugging","Code optimization"]

Design, build, and optimize GPU-accelerated deep learning inference software for large-scale LLM and generative AI serving. Contribute to high-performance, open-source inference frameworks and NVIDIA libraries, driving performance improvements across NVIDIA accelerators from datacenter GPUs to edge SoCs. Implement and tune model serving pipelines using CUDA kernels and tools such as CUTLASS, OAI Triton, and NCCL, collaborating across cross-functional teams to ship impactful features for real deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Senior Deep Learning Software Engineer, Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design, build, and optimize GPU-accelerated deep learning inference software for large-scale LLM and generative AI serving. Contribute to high-performance, open-source inference frameworks and NVIDIA libraries, driving performance improvements across NVIDIA accelerators from datacenter GPUs to edge SoCs. Implement and tune model serving pipelines using CUDA kernels and tools such as CUTLASS, OAI Triton, and NCCL, collaborating across cross-functional teams to ship impactful features for real deployments.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Performance optimization, analysis, and tuning of deep learning models across domains like LLM, multimodal, and generative AI.
  • •Scale deep learning inference performance across architectures and NVIDIA accelerator types, from datacenter GPUs to edge SoCs.
  • •Contribute features and code to inference libraries and solutions including vLLM, SGLang, FlashInfer, and other LLM software.
  • •Implement and optimize model serving pipelines using open-source tools and plugins such as CUTLASS, OAI Triton, NCCL, and CUDA kernels.
  • •Collaborate with cross-functional teams across frameworks and NVIDIA inference optimization efforts to improve platform performance and deployment.

Pay and Benefits

Salary: USD 152,000 - 287,500
Equity and Bonus:Equity

Key Requirements

  • •Masters or PhD (or equivalent) in a relevant field such as Computer Engineering, Computer Science, EECS, or AI.
  • •5+ years of relevant software development experience.
  • •Excellent C/C++ programming and software design skills; SW Agile skills helpful and Python experience a plus.
  • •Prior experience training, deploying, or optimizing inference of deep learning models in production is a plus.
  • •Experience with performance modeling, profiling, debugging, and code optimization or architectural knowledge of CPU and GPU is a plus.
Experience:5+ yearsDeep learningLLMGenerative AIGPU accelerationOpen source
Education:
Skills:Performance optimizationAnalysisTuningSoftware designCross-collaborationDebuggingCode optimization
Tech Stack:CC++PythonCUTLASSOAI TritonNCCLCUDAVLLMSGLangFlashInferPyTorchNVSHMEMCPUGPULLMMultimodalGenerative AI

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor