Senior Staff SRE – Compute Platform

NVIDIA
Bengaluru
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsSkills: ["Communication","Incident management","Cross-team collaboration","Problem-solving","Delivering scalable solutions"]

Build and operate large-scale Kubernetes and KubeVirt compute platforms that power global engineering workloads. Own bare-metal provisioning and lifecycle management across data centers, then automate provisioning, self-service, and observability using APIs and infrastructure tooling. Drive reliability by defining SLOs/SLIs, managing incidents, and leading postmortems. Partner with infrastructure, security, hardware, and application teams, and support an on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Senior Staff SRE – Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 28 minutes agoStatus: Live

Job Summary

Build and operate large-scale Kubernetes and KubeVirt compute platforms that power global engineering workloads. Own bare-metal provisioning and lifecycle management across data centers, then automate provisioning, self-service, and observability using APIs and infrastructure tooling. Drive reliability by defining SLOs/SLIs, managing incidents, and leading postmortems. Partner with infrastructure, security, hardware, and application teams, and support an on-call rotation.
Location: Bengaluru
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, operate, and improve large-scale Kubernetes and KubeVirt compute platforms with a focus on performance, capacity, reliability, and operational scale.
  • •Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, networking services, OS provisioning, hardware validation, and fleet automation.
  • •Develop automation, self-service, and observability capabilities using APIs, Python or Go, and Infrastructure as Code and configuration management.
  • •Define and operate SLOs/SLIs/error budgets, manage alerting and incident response, and lead incident investigations and blameless postmortems.
  • •Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives and participate in an on-call rotation.

Key Requirements

  • •BS in Computer Science/Engineering or equivalent experience, plus 10+ years operating production infrastructure or platform services.
  • •Strong Kubernetes administration and KubeVirt experience, including containerization and resolving distributed-system challenges on Linux.
  • •Experience deploying and operating bare-metal infrastructure in data centers, including provisioning, networking, OS lifecycle management, and hardware automation.
  • •Proficiency in Python or Go (or comparable), building RESTful services and integrating infrastructure APIs.
  • •Experience with Infrastructure as Code/automation (e.g., Terraform, Ansible, Chef, Puppet) and SRE observability/tooling (SLIs/SLOs, incident management, OpenTelemetry, Prometheus, Grafana, ELK Stack, Splunk).
Experience:10+ yearsInfrastructurePlatform servicesKubernetesBare-metalSREObservabilityData centers
Education:
Skills:CommunicationIncident managementCross-team collaborationProblem-solvingDelivering scalable solutions
Tech Stack:KubernetesKubeVirtLinuxDockerContainerizationMicroservicesPXE bootDHCPDNSPythonGoRESTful servicesInfrastructure as CodeTerraformAnsibleChefPuppetTCP/IPOpenTelemetryPrometheus

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor