Senior Solutions Architect, HPC and AI

NVIDIA
Germany, Switzerland, France
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsSkills: ["Problem-solving","Performance optimization","Collaboration","Debugging","Troubleshooting"]

Deploy, debug, and optimize AI training and inference workloads on large-scale GPU clusters, partnering with internal framework developers and external customers across Europe. Benchmark new framework capabilities, analyze performance, and share actionable insights. Help customers resolve cluster stability and bottleneck issues, scale reliably on NVIDIA GPUs, and contribute to Europe’s Sovereign AI initiative by improving resiliency features in training pipelines.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior Solutions Architect, HPC and AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Deploy, debug, and optimize AI training and inference workloads on large-scale GPU clusters, partnering with internal framework developers and external customers across Europe. Benchmark new framework capabilities, analyze performance, and share actionable insights. Help customers resolve cluster stability and bottleneck issues, scale reliably on NVIDIA GPUs, and contribute to Europe’s Sovereign AI initiative by improving resiliency features in training pipelines.
Location: Germany, Switzerland, France
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Collaborate with training framework developers and product teams to stay ahead of new features and help partners adopt them effectively.
  • •Support deployment and debugging of AI workloads, improving efficiency on NVIDIA platforms.
  • •Benchmark framework features, analyze performance, and share actionable insights with customers and internal teams.
  • •Work directly with customers to solve cluster performance/stability issues, identify bottlenecks, and implement solutions.
  • •Guide customers in scaling workloads efficiently and reliably on latest-generation NVIDIA GPUs, including contributing to resiliency features in AI training pipelines.

Key Requirements

  • •BS, MS, PhD (or equivalent) in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering field (or equivalent practical experience).
  • •8+ years of experience in accelerated computing at cluster scale, ideally with NVIDIA platforms.
  • •Strong programming skills in C, C++, or Python.
  • •Hands-on experience profiling and debugging large parallel applications and resolving bottlenecks in large-scale training workloads or parallel applications.
  • •Solid understanding of CPU/GPU architectures, CUDA, parallel filesystems, high-speed interconnects, and cluster scheduling/resource management (e.g., SLURM or cloud clusters).
Experience:8+ yearsAccelerated computingHPCAIGPU clustersTraining/inference
Education:
Skills:Problem-solvingPerformance optimizationCollaborationDebuggingTroubleshooting
Tech Stack:CUDACC++PythonNsight SystemsNsight ComputeNCCLMPISLURMPyTorchMegatron-LMNeMoVLLMDynamoTensorRT-LLMRedHat Inference ServerSGLangParallel filesystemsHigh-speed interconnectsLLM frameworks

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor