Senior Staff Site Reliability Engineer – Compute Platform

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Communication","Interpersonal communication","Incident management","Problem-solving"]

Build and operate reliable, scalable compute platforms powering global engineering workloads. Own large-scale Kubernetes and KubeVirt environments, bare-metal provisioning and lifecycle management, and fleet automation. Develop self-service, automation, and observability using APIs, Python or Go, Infrastructure as Code, and metrics/logs/traces. Define and run SLOs/SLIs with strong incident-response practices while partnering across infrastructure, security, hardware, and applications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 hours ago

Senior Staff Site Reliability Engineer – Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and operate reliable, scalable compute platforms powering global engineering workloads. Own large-scale Kubernetes and KubeVirt environments, bare-metal provisioning and lifecycle management, and fleet automation. Develop self-service, automation, and observability using APIs, Python or Go, Infrastructure as Code, and metrics/logs/traces. Define and run SLOs/SLIs with strong incident-response practices while partnering across infrastructure, security, hardware, and applications.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms focused on performance, capacity, reliability, and operational scale.
  • •Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.
  • •Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, and Infrastructure as Code with metrics/logs/traces and service-health data.
  • •Define and operate SLOs/SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations and blameless postmortems.
  • •Partner with infrastructure, security, hardware, data-center, and application teams on global platform initiatives and participate in an on-call rotation.

Key Requirements

  • •BS in Computer Science/Engineering or related field (or equivalent experience), plus 10+ years operating production infrastructure or platform services.
  • •Strong Kubernetes administration expertise, including KubeVirt, Docker/containerization, microservices, and Linux for distributed-systems troubleshooting.
  • •Experience deploying and operating bare-metal infrastructure in data-center environments, including provisioning, networking, OS lifecycle management, and hardware automation.
  • •Proficiency in Python, Go, or a comparable language; experience building RESTful services and integrating infrastructure APIs.
  • •Infrastructure as Code and automation experience (Terraform, Ansible, Chef, or Puppet) plus SRE/observability skills with SLIs/SLOs, incident management, and tools like OpenTelemetry/Prometheus/Grafana/ELK/Splunk.
Experience:10+ yearsSREKubernetesBare-metal infrastructureInfrastructure automationObservabilityDistributed systems
Education:Bachelor's in Computer Science, Engineering (or related technical field)
Skills:CommunicationInterpersonal communicationIncident managementProblem-solving
Tech Stack:KubernetesKubeVirtLinuxDockerPythonGoRESTful servicesInfrastructure as CodeTerraformAnsibleChefPuppetTCP/IPOpenTelemetryPrometheusGrafanaELK StackSplunkPXE bootDHCP

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor