Service Reliability Engineer

NVIDIA
Texas
Workplace: RemoteFull timeUSD 168,000 - 333,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Documentation","Mentoring","Customer focus"]

As an SRE on the Compute Infrastructure Support team, you’ll run 24/7 follow-the-sun support and help maintain near-100% availability for on-prem and cloud AI infrastructure. Monitor and manage production GPU and Kubernetes environments, use observability to diagnose incidents, and build predictive automation and auto-healing solutions. Collaborate with service owners to resolve complex issues and continuously improve operational processes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
22 hours ago

Service Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

As an SRE on the Compute Infrastructure Support team, you’ll run 24/7 follow-the-sun support and help maintain near-100% availability for on-prem and cloud AI infrastructure. Monitor and manage production GPU and Kubernetes environments, use observability to diagnose incidents, and build predictive automation and auto-healing solutions. Collaborate with service owners to resolve complex issues and continuously improve operational processes.
Location: Texas
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Operate within a 24/7 follow-the-sun support model with global coverage and manage scheduled shifts.
  • •Monitor and manage extensive production GPU and Kubernetes environments to ensure high availability and performance.
  • •Detect, prevent, and respond to incidents using advanced tools; diagnose issues via logs, metrics, and system behavior and implement resolutions.
  • •Develop predictive automated support routines and improve automation with auto-healing and automated break-fix solutions.
  • •Perform systems administration, network administration, and security monitoring; coordinate with domain experts and service owners to resolve complex issues.

Pay and Benefits

Salary: USD 168,000 - 333,500 annually
Equity and Bonus:Equity

Key Requirements

  • •8+ years coordinating large-scale production systems and 3+ years in high-availability Internet, Cloud, or Data Center environments.
  • •Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
  • •Expert-level Linux system administration and automation using Ansible and/or Python, with strong shell scripting and networking fundamentals.
  • •Proficiency with observability and incident management tools such as Grafana, OpenTelemetry, PagerDuty, and JIRA.
  • •BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience.
Experience:8+ yearsHigh-availabilityCloudData center
Education:Bachelor's
Skills:CommunicationProblem-solvingDocumentationMentoringCustomer focus
Tech Stack:KubernetesSLURMGrafanaOpenTelemetryPagerDutyJIRAAWSAzureGCPOCILinuxAnsiblePythonShell scriptingDNSDHCPStorage systemsNetworkingGPUCloud platforms

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor