HPC Engineer

NVIDIA
Singapore
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsEducation: bachelorsSkills: ["Interpersonal skills","Verbal and written communication","Organizational skills","Problem-solving","Customer-focused","Ability to troubleshoot"]

Deploy, manage, and maintain AI/HPC infrastructure in Linux-based environments for new and existing customers. Serve as a domain expert during planning calls, implementing large-scale AI/HPC projects and supporting customer rollouts through documentation and knowledge transfers. Triage customer issues, provide feedback to internal teams, and contribute to improvements via bug reports and workarounds. Work across networking, system design, and automation to deliver large-scale systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

HPC Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Deploy, manage, and maintain AI/HPC infrastructure in Linux-based environments for new and existing customers. Serve as a domain expert during planning calls, implementing large-scale AI/HPC projects and supporting customer rollouts through documentation and knowledge transfers. Triage customer issues, provide feedback to internal teams, and contribute to improvements via bug reports and workarounds. Work across networking, system design, and automation to deliver large-scale systems.
Location: Singapore
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Deploy, manage, and maintain AI/HPC infrastructure in Linux-based environments for new and existing customers.
  • •Act as a domain expert with customers during planning calls and through implementation of large-scale AI/HPC projects.
  • •Prepare handover documentation and perform knowledge transfers to support customer rollouts.
  • •Provide feedback to internal teams by opening bugs, documenting workarounds, and suggesting improvements.
  • •Maintain and deliver resolutions for customer-blocking issues as they arise.

Key Requirements

  • •Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • •3+ years providing in-depth support and deployment services, solving problems for hardware and software products.
  • •Knowledge and experience with AI Factory / HPC and Linux system administration, including process management, task scheduling, kernel management, boot procedures, troubleshooting, performance reporting, and optimization.
  • •Experience with cluster management technologies and scripting.
  • •Experience with AI/HPC schedulers such as SLURM, LSF, and PBS (or similar).
Experience:3+ yearsAI/HPCHPCInfrastructure
Education:Bachelor's
Skills:Interpersonal skillsVerbal and written communicationOrganizational skillsProblem-solvingCustomer-focusedAbility to troubleshoot
Certifications:Linux certifications
Languages:English
Tech Stack:LinuxAI/HPCAI FactoryCluster managementProcess managementTask schedulingKernel managementBoot proceduresTroubleshootingPerformance reportingOptimizationSLURMLSFPBSScriptingInfiniBandEthernetGPUMPIAnsible

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor