NIM Solution Architect

NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 2+ yearsEducation: mastersSkills: ["Written communication","Verbal communication","Collaboration"]

Work hands-on on NVIDIA Inference Microservices (NIM) to implement, deploy, and optimize inference solutions for enterprise and industry AI workloads. Package and serve open-source, NVIDIA, and customer models via standardized, containerized NIM APIs across on-prem, cloud, and hybrid environments. Tune NIM models, optimize high-volume LLM/VLM inference and rollout workloads, and deliver demos and client support. Partner with cross-functional teams to expand the AI solutions portfolio.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
13 hours ago

NIM Solution Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Work hands-on on NVIDIA Inference Microservices (NIM) to implement, deploy, and optimize inference solutions for enterprise and industry AI workloads. Package and serve open-source, NVIDIA, and customer models via standardized, containerized NIM APIs across on-prem, cloud, and hybrid environments. Tune NIM models, optimize high-volume LLM/VLM inference and rollout workloads, and deliver demos and client support. Partner with cross-functional teams to expand the AI solutions portfolio.
Location: Shanghai, Beijing
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Drive implementation, deployment, and optimization of NVIDIA Inference Microservices (NIM) solutions for enterprise and industry AI workloads.
  • •Package and serve open-source, NVIDIA, and customer-proprietary models through NIM using standardized, containerized APIs for on-premises, cloud, and hybrid environments.
  • •Optimize high-volume inference and rollout workloads for LLMs and VLMs, including evaluating and tuning NIM models.
  • •Deliver technical projects, demos, and client support tasks as directed by Solution Architecture leadership.
  • •Provide technical support and guidance to customers to facilitate adoption and implementation of NVIDIA technologies and products.

Key Requirements

  • •Master’s degree or higher in Computer Science, Machine Learning, Electrical Engineering, Mathematics, or a related technical field, or equivalent experience.
  • •2+ years of hands-on experience in machine learning engineering, applied research, LLM/VLM inference, or RL rollout.
  • •Production-quality Python and PyTorch skills, including distributed GPU training, profiling, debugging, and memory optimization.
  • •Working knowledge of transformer architectures, performance optimization, rollout sampling strategies, structured generation, and model-quality evaluation.
  • •Strong written and verbal communication skills to collaborate across research, engineering, infrastructure, product, and customer-facing teams.
Experience:2+ yearsMachine learningApplied researchLLM/VLM inferenceRLGenerative AIEnterprise AI deployment
Education:Master's
Skills:Written communicationVerbal communicationCollaboration
Tech Stack:PythonPyTorchDistributed GPU trainingTransformer architecturesContainerized APIsAPIsProfilingDebuggingMemory optimizationSLIMENemo-RL

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor