Senior HPC Platform Architect

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Problem-solving","Communication"]

Own data center architecture reviews for new HPC clusters, evaluating compute, storage, networking, and cooling decisions across GPU compute, AI infrastructure, and EDA environments. Serve as the BDC representative to drive architecture alignment, validate topology and latency tradeoffs, and lead performance benchmarking, profiling, and cluster health checks. Tune scheduler, hardware, and OS/kernel parameters (NUMA, huge pages, kernel images) and partner with vendors and teams to optimize and improve HPC observability, tooling, and documentation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 days ago

Senior HPC Platform Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 28 minutes agoStatus: Live

Job Summary

Own data center architecture reviews for new HPC clusters, evaluating compute, storage, networking, and cooling decisions across GPU compute, AI infrastructure, and EDA environments. Serve as the BDC representative to drive architecture alignment, validate topology and latency tradeoffs, and lead performance benchmarking, profiling, and cluster health checks. Tune scheduler, hardware, and OS/kernel parameters (NUMA, huge pages, kernel images) and partner with vendors and teams to optimize and improve HPC observability, tooling, and documentation.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own data center architecture reviews for new HPC clusters, evaluating compute, storage, networking, and cooling decisions and challenging assumptions before buildout.
  • •Represent BDC in cluster build meetings across simultaneous programs (GPU compute clusters, AI infrastructure, and EDA environments) to drive architecture alignment.
  • •Analyze and validate cluster design choices across storage-to-compute distance, cross-mount latency, rack layout, and multi-site topology, surfacing risks and tradeoff recommendations to leadership.
  • •Lead performance benchmarking and profiling, running sanity benchmarks, regression suites, and cluster health checks to identify bottlenecks early.
  • •Drive infrastructure optimization across scheduler-level tuning (LSF/Slurm), hardware-level tuning (GPU, CPU, networking, storage), and OS/kernel-level tuning (NUMA binding, socket binding, huge pages, kernel image selection).

Key Requirements

  • •B.E./B.Tech or M.Tech/M.S. with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale.
  • •Deep understanding of data center architecture fundamentals across compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand, Ethernet, NVLink).
  • •Proven ability to evaluate and challenge infrastructure design decisions including rack layout, power/cooling constraints, storage-to-compute distance, and network fabric topology.
  • •Experience with OS and kernel-level performance tuning such as NUMA binding, socket affinity, huge page configuration, kernel image selection, and system parameter optimization.
  • •Hands-on experience administering large-scale Linux HPC clusters using workload managers such as LSF and/or Slurm, plus scripting ability (Python, Bash, or Perl) for automation and analysis.
Experience:5+ yearsHPCData centerLinuxSREPlatform engineering
Skills:Problem-solvingCommunication
Tech Stack:HPC infrastructureData center architectureLinuxUnixPythonBashPerlLSFSlurmNUMAKernel imageHuge page configurationInfiniBandEthernetNVLink

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor