Software Engineer, LLM Inference

NVIDIA
Shanghai
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 4+ yearsEducation: mastersSkills: ["Communication","Customer communication","Collaboration","Proactivity","Debugging"]

Develop and scale robust LLM inference software across multiple platforms, focusing on functionality, performance, and tuning. Conduct performance analysis and optimization, and stay current with AI/LLM research to improve model inferencing. Work on TensorRT and TensorRT Edge LLM updates, and collaborate with software, research, and product teams to guide machine-learning inferencing direction.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
10 hours ago

Software Engineer, LLM Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live
Reposted: similar role first listed 8 months ago

Job Summary

Develop and scale robust LLM inference software across multiple platforms, focusing on functionality, performance, and tuning. Conduct performance analysis and optimization, and stay current with AI/LLM research to improve model inferencing. Work on TensorRT and TensorRT Edge LLM updates, and collaborate with software, research, and product teams to guide machine-learning inferencing direction.
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Craft and develop robust inferencing software that can be scaled to multiple platforms.
  • •Perform performance analysis, optimization, and tuning for inference workloads.
  • •Follow academic developments in AI and incorporate feature updates using TensorRT and TensorRT Edge LLM.
  • •Collaborate across software, research, and product teams to guide machine learning inferencing direction.
  • •Provide highly responsive support when needed, including customer communication.

Key Requirements

  • •Masters or higher degree in Computer Engineering, Computer Science, Applied Mathematics, or a related computing-focused field (or equivalent experience).
  • •4+ years of relevant software development experience.
  • •Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
  • •Experience working with deep learning frameworks like PyTorch.
  • •Excellent written and oral communication skills in English.
Experience:4+ years
Education:Master's
Skills:CommunicationCustomer communicationCollaborationProactivityDebugging
Languages:English
Tech Stack:CC++CUDACuDNNTensorRTTensorRT EdgePyTorchLLMsDeep learning

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor