Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 224,000 - 356,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 12+ yearsEducation: bachelorsSkills: ["Troubleshooting","Judgment","Communication","Prioritization","Time management"]

Raise the Day 2 operations bar across the NVIDIA Cloud Partner ecosystem by solving large-scale reliability, performance, stability, efficiency, and cost problems after cluster validation. Partner with NCP engineers to prototype and validate operational approaches, prepare operating models for new NVIDIA platforms, and close maturity gaps across people, process, tooling, telemetry, security, and incident response—turning proven work into repeatable ecosystem capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
36 minutes ago

Senior Solutions Architect, NVIDIA Cloud Partner Operations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 36 minutes agoStatus: Live

Job Summary

Raise the Day 2 operations bar across the NVIDIA Cloud Partner ecosystem by solving large-scale reliability, performance, stability, efficiency, and cost problems after cluster validation. Partner with NCP engineers to prototype and validate operational approaches, prepare operating models for new NVIDIA platforms, and close maturity gaps across people, process, tooling, telemetry, security, and incident response—turning proven work into repeatable ecosystem capabilities.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Solve hard Day 2 operations problems at scale by working with partner engineers to identify causes, prototype approaches, validate under representative load, and leave behind an operational practice.
  • •Prepare new NVIDIA technology for Day 2 by helping partners set up operating models for new platforms, capacity, services, and use cases before customer dependence and during live adoption.
  • •Improve reliability, performance, and economics together using incident frequency, recovery time, utilization, and cost-per-token to measure losses and validate fixes.
  • •Raise each partner’s Day 2 maturity by identifying and closing gaps across people, process, tooling, telemetry, security, and incident response.
  • •Turn validated solutions into ecosystem capability by creating operating procedures, reference architectures, assessments, automation, and agentic workflows for other NCPs.

Pay and Benefits

Salary: USD 224,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field (or equivalent experience).
  • •12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar role; or 5+ years of exceptional specialist-level work in large-scale GPU/AI infrastructure.
  • •Experience building, operating, or improving distributed infrastructure under real production load.
  • •Deep expertise in at least one part of the Day 2 stack with hands-on work involving large-scale GPU, HPC, or cloud infrastructure (e.g., DCGM, BMC/Redfish, InfiniBand/high-speed Ethernet, NCCL/UFM, or high-performance storage).
  • •Hands-on experience with operating platforms such as Kubernetes or Slurm; observability with Prometheus/Grafana or OpenTelemetry; and automation with Terraform/Ansible/Argo CD, plus strong Linux knowledge and enough Python/Bash to automate diagnosis and remediation.
Experience:12+ yearsAI infrastructureGPU cloudHPCCloud engineeringDistributed infrastructureSite reliability engineering
Education:Bachelor's in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or related field
Skills:TroubleshootingJudgmentCommunicationPrioritizationTime management
Tech Stack:DCGMBMC/RedfishFirmwareDriver lifecycleInfiniBandHigh-speed EthernetNCCLUFMLustreIBM Storage ScaleWEKAVAST DataKubernetesSlurmGPU schedulingMulti-tenancyPrometheusGrafanaOpenTelemetryTerraform

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor