Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA
Santa Clara, Texas, Virginia, Washington
Workplace: HybridFull timeUSD 148,000 - 235,750 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 6+ yearsEducation: bachelorsSkills: ["Communication","Cross-functional collaboration"]

Drive first-time deployment validation for NVIDIA’s AI Factories, from first rack power-on through customer handoff. Run and debug AI/LLM workloads and performance benchmarks on Linux-based multi-GPU, multi-node clusters using NCCL and collectives like AllReduce and AllToAll. Operationalize observability and automation to capture structured evidence, troubleshoot failures, and recommend configuration changes to improve throughput, latency, and scaling efficiency across cross-functional teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Senior Solutions Architect, First Time Deployment Validation - NVIS

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Drive first-time deployment validation for NVIDIA’s AI Factories, from first rack power-on through customer handoff. Run and debug AI/LLM workloads and performance benchmarks on Linux-based multi-GPU, multi-node clusters using NCCL and collectives like AllReduce and AllToAll. Operationalize observability and automation to capture structured evidence, troubleshoot failures, and recommend configuration changes to improve throughput, latency, and scaling efficiency across cross-functional teams.
Location: Santa Clara, Texas, Virginia, Washington
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • •Ensure configurations align with NCCL, collectives, and distributed training frameworks.
  • •Own AI/LLM benchmark execution: setup, orchestration, result collection, and analysis.
  • •Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.
  • •Operationalize observability and automation (Python/Shell) to capture structured evidence and drive validation readiness.

Pay and Benefits

Salary: USD 148,000 - 235,750 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or a related field.
  • •More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
  • •Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters with practical NCCL knowledge.
  • •Knowledge of collective communication patterns, especially AllReduce and AllToAll, for contemporary ML/LLM training.
  • •Proficiency in Python and Shell/Bash for scripting, automation, tooling, and benchmarking.
Experience:6+ yearsHPCDistributed systemsAI/MLLLMObservability
Education:Bachelor's
Skills:CommunicationCross-functional collaboration
Tech Stack:LinuxGPU clustersNCCLAllReduceAllToAllPythonShell/BashPyTorchTensorFlowMetricsLogsTracesDashboards

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor