Software Platform Support Engineer - GPU Cloud

NVIDIA
Santa Clara, Durham, North Carolina, Virginia, Washington
Full timeUSD 108,000 - 172,500 annuallyFunction: Product ManagementExperience: 5+ yearsEducation: high_schoolSkills: ["Customer service","Troubleshooting","Communication","Organizational skills"]

Support internal customers using NVIDIA DGX Cloud platforms by understanding their workloads, troubleshooting complex cloud deployments, and documenting solutions for faster self-service. Coordinate with internal teams for Tier 1 support, triage and escalate root-cause issues, and file bugs with the Site Reliability team. Improve operational runbooks and workflows, build tooling to increase support visibility, and participate in an on-call rotation across distributed, compute/storage/network environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Software Platform Support Engineer - GPU Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Support internal customers using NVIDIA DGX Cloud platforms by understanding their workloads, troubleshooting complex cloud deployments, and documenting solutions for faster self-service. Coordinate with internal teams for Tier 1 support, triage and escalate root-cause issues, and file bugs with the Site Reliability team. Improve operational runbooks and workflows, build tooling to increase support visibility, and participate in an on-call rotation across distributed, compute/storage/network environments.
Location: Santa Clara, Durham, North Carolina, Virginia, Washington
Employment Type: Full time
Job Function: Product Management
Seniority: Mid level

Key Responsibilities

  • •Coordinate with multiple internal teams to provide Tier 1 support for complex cloud platforms.
  • •Define and improve operational workflows such as runbooks, escalation paths, and support processes.
  • •Triage and investigate root cause of customer issues, escalating as needed.
  • •File bugs and report issues while working closely with the Site Reliability team.
  • •Build tooling to improve customer support process and visibility, including on-call support for production systems.

Pay and Benefits

Salary: USD 108,000 - 172,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS/MS degree in Computer Science or related areas (or equivalent experience).
  • •5+ years supporting distributed software systems and end-user software platforms, with Linux experience.
  • •Hands-on experience with Kubernetes and major cloud platforms including AWS, Azure, OCI, and GCP.
  • •Background in infrastructure, networking, storage, and DevOps scripting/tooling.
  • •Customer support experience with strong troubleshooting and communication skills.
Experience:5+ yearsCloud deploymentsDistributed systemsDevOpsSRE
Education:High School in Computer science or related areas
Skills:Customer serviceTroubleshootingCommunicationOrganizational skills
Tech Stack:LinuxKubernetesAWSAzureOCIGCPDevOpsRunbooksEscalation pathsData storage technologiesDatabasesFileBlockBlobToolingSLURMHPCMLOpsGPU workloadsDistributed training systems

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor