Principal Engineer, Service Delivery

Dell
Cyberjaya
Workplace: RemoteFull timeFunction: Transportation & Fleet OperationsExperience: 8+ yearsEducation: bachelorsSkills: ["Troubleshooting","Problem-solving","Stakeholder management"]

Serve as a Senior GenAI & HPC Engineer on the Service Delivery team, supporting onsite deployments across South East Asia/APJ. Build, integrate, and test large multi-GPU systems, benchmark them with industry-standard tools, and recommend performance optimizations. Design and deliver advanced service solutions using GPU-accelerated compute clusters (NVIDIA), Linux performance tuning, and large-scale GPU networking (including InfiniBand and L2/L3 leaf-spine configurations).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Dell
Dell
1 day ago

Principal Engineer, Service Delivery

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Serve as a Senior GenAI & HPC Engineer on the Service Delivery team, supporting onsite deployments across South East Asia/APJ. Build, integrate, and test large multi-GPU systems, benchmark them with industry-standard tools, and recommend performance optimizations. Design and deliver advanced service solutions using GPU-accelerated compute clusters (NVIDIA), Linux performance tuning, and large-scale GPU networking (including InfiniBand and L2/L3 leaf-spine configurations).
Location: Cyberjaya
Workplace: Remote
Employment Type: Full time
Job Function: Transportation & Fleet Operations
Seniority: Mid level

Key Responsibilities

  • •Design and deliver advanced service solutions for advanced HPC and GenAI deployments.
  • •Build, integrate, and test large multi-GPU systems and benchmark them using industry-standard tools.
  • •Perform performance optimization by analyzing results and making recommendations for GPU and infrastructure tuning.
  • •Conduct root-cause analyses and corrective actions for operational issues.
  • •Develop automation and monitoring to reduce toil and prepare handover/operational documentation.
Travel: High travel

Key Requirements

  • •8+ years of related experience.
  • •Experience deploying GPU-accelerated AI compute clusters for NVIDIA Base Command Manager in an NVL72 environment.
  • •Experience implementing large-scale GPU networking (more than 100,000 connections).
  • •Hands-on networking configuration using Nvidia Spectrum switches/Cumulus OS and InfiniBand switches.
  • •Strong troubleshooting and stakeholder management skills; network cabling design and air/liquid-cooled racks are added experience.
Experience:8+ yearsHPCGenAIGPU clustersGPU networking
Education:Bachelor's in Engineering, Computer Science, or related field
Skills:TroubleshootingProblem-solvingStakeholder management
Tech Stack:GPUNVIDIA Base Command ManagerNVL72Nvidia Spectrum switchesCumulus OSInfiniBand switchesSonic OSKubernetesUbuntuOpenShiftAir CooledLiquid cooled

Company Brief

Dell
Designs, manufactures and sells computers, servers, storage, networking equipment, and related IT solutions and services for consumers, small businesses, and enterprises worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Round Rock, United States
Founded: 1984
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor