Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA
Santa Clara, United States
Workplace: RemoteFull timeUSD 224,000 - 356,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Troubleshooting","Communication","Prioritization","Time management","Leading cross-functional work"]

Work hands-on with NVIDIA Cloud Partners to elevate Day 2 operations across GPU cloud and AI infrastructure. You’ll solve large-scale reliability, performance, stability, and efficiency problems, make new technologies production-ready, and raise partners’ Day 2 maturity. Turn validated work into reusable operating procedures and reference architectures, and build feedback loops that help account teams, support, product, and engineering fix recurring issues at the right level.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Solutions Architect, NVIDIA Cloud Partner Operations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Work hands-on with NVIDIA Cloud Partners to elevate Day 2 operations across GPU cloud and AI infrastructure. You’ll solve large-scale reliability, performance, stability, and efficiency problems, make new technologies production-ready, and raise partners’ Day 2 maturity. Turn validated work into reusable operating procedures and reference architectures, and build feedback loops that help account teams, support, product, and engineering fix recurring issues at the right level.
Location: Santa Clara, United States
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Solve Day 2 operations problems at scale by working with partner engineers to diagnose causes, prototype approaches, validate under representative load, and leave behind an operable practice.
  • •Prepare and adapt new NVIDIA platforms and partner operating models for production readiness, improving adoption in live environments without degrading service.
  • •Improve reliability, performance, and economics using metrics such as incident frequency, recovery time, utilization, and cost per token to measure gaps and validate fixes.
  • •Raise each partner’s Day 2 maturity across people, process, tooling, telemetry, security, and incident response.
  • •Convert validated work into ecosystem capability via operating procedures, reference architectures, assessments, automation, and agentic workflows that other partners can integrate.

Pay and Benefits

Salary: USD 224,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related field (or equivalent experience).
  • •12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or similar; or 5+ years of exceptional specialist-level work in large-scale GPU/AI infrastructure.
  • •Experience building, operating, or improving distributed infrastructure under real production load.
  • •Deep expertise in a Day 2 stack area with hands-on work (examples include DCGM, BMC/Redfish, firmware/driver lifecycle; InfiniBand/high-speed Ethernet, NCCL, UFM; or high-performance storage like Lustre, IBM Storage Scale, WEKA, or VAST Data).
  • •Hands-on experience across the operating platform including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus/Grafana/OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar.
Experience:5+ yearsCloud infrastructureAI infrastructureHPCGPU cloudSite reliability engineeringDistributed systems
Education:Bachelor's
Skills:TroubleshootingCommunicationPrioritizationTime managementLeading cross-functional work
Tech Stack:Day 2 operationsDCGMBMC/RedfishInfiniBandHigh-speed EthernetNCCLUFMLustreIBM Storage ScaleWEKAVAST DataKubernetesSlurmGPU schedulingMulti-tenancyPrometheusGrafanaOpenTelemetryTerraformAnsible

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor