AI Computing Software Development Engineer, TensorRT-LLM

NVIDIA
Taipei
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsEducation: mastersSkills: ["Proactive","Independent","Communication","Team collaboration","Problem-solving"]

Develop and optimize scalable LLM inference software for NVIDIA’s TensorRT-LLM team. You will analyze and improve LLM inference performance, implement kernels and runtime features for new models and algorithms, and stay current with academic and industrial advances. Collaborate across software, research, and product groups to shape deep learning inference architecture and hardware direction, while ensuring high-quality, high-performance implementation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
22 hours ago

AI Computing Software Development Engineer, TensorRT-LLM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live
Reposted: similar role first listed 7 months ago

Job Summary

Develop and optimize scalable LLM inference software for NVIDIA’s TensorRT-LLM team. You will analyze and improve LLM inference performance, implement kernels and runtime features for new models and algorithms, and stay current with academic and industrial advances. Collaborate across software, research, and product groups to shape deep learning inference architecture and hardware direction, while ensuring high-quality, high-performance implementation.
Location: Taipei
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Craft and develop robust inference software scaled across multiple platforms for functionality and performance.
  • •Perform performance analysis and optimization for LLM inference.
  • •Follow academic and industrial developments in AI and feature-update TensorRT-LLM.
  • •Implement kernels and runtime features to support new LLM models and inference algorithms.
  • •Provide feedback into inference architecture and hardware design and development.

Key Requirements

  • •Master’s degree or higher in Computer Engineering, Computer Science, Applied Mathematics, or a related computing-focused field (or equivalent experience).
  • •3+ years of relevant software development experience.
  • •Excellent Python programming, software design, and software engineering skills.
  • •Experience working with deep learning frameworks like PyTorch and HuggingFace.
  • •Awareness of the latest developments in LLM architectures and LLM inference techniques.
Experience:3+ yearsDeep learningLLM inferenceHPC
Education:Master's
Skills:ProactiveIndependentCommunicationTeam collaborationProblem-solving
Languages:English
Tech Stack:PythonPyTorchHuggingFaceTensorRT-LLMCC++CUDAOpenCLLLM inferenceDeep learningKernelsRuntime featuresPerformance analysisProfilingDebuggingTest design

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor