Site Reliability Engineer

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Curiosity","Communication","Teamwork","Ownership"]

Support SRE initiatives to improve reliability, scalability, and developer efficiency across NVIDIA enterprise systems. Build and maintain distributed, cloud-native services, automate database operations (provisioning, scaling, backup, failover), and enhance observability with dashboards, alerts, and automation. Participate in incident response to reduce MTTR and drive post-incident improvements. Collaborate with Cloud, Platform, Security, and AI/ML teams while operating Kubernetes-based infrastructure and adopting AI-assisted engineering practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Support SRE initiatives to improve reliability, scalability, and developer efficiency across NVIDIA enterprise systems. Build and maintain distributed, cloud-native services, automate database operations (provisioning, scaling, backup, failover), and enhance observability with dashboards, alerts, and automation. Participate in incident response to reduce MTTR and drive post-incident improvements. Collaborate with Cloud, Platform, Security, and AI/ML teams while operating Kubernetes-based infrastructure and adopting AI-assisted engineering practices.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Support SRE initiatives to improve reliability, scalability, and developer efficiency across enterprise systems.
  • •Build and maintain distributed systems that power NVIDIA AI-powered enterprise products and services.
  • •Automate database operations, including provisioning, scaling, backup, and failover for relational and vector database services.
  • •Improve observability with dashboards, alerts, and automation scripts; participate in incident response and post-incident reviews.
  • •Collaborate with Cloud, Platform, Security, and AI/ML teams to support platform reliability, operate Kubernetes-based infrastructure, and apply SRE best practices.

Key Requirements

  • •BS degree in Computer Science or a related technical field, or equivalent practical experience.
  • •Foundational proficiency in at least one programming language such as Python, TypeScript, JavaScript, or Go.
  • •Basic understanding of cloud platforms (AWS, Azure, or GCP) and container technologies like Docker and Kubernetes.
  • •Exposure to infrastructure-as-code tools (e.g., Terraform, AWS CDK, CloudFormation) or willingness to learn.
  • •Familiarity with Linux/Unix, networking fundamentals, version control (Git), and observability concepts/tools.
Experience:DevOpsSRECloud infrastructureAutomation
Education:Bachelor's
Skills:Problem-solvingCuriosityCommunicationTeamworkOwnership
Tech Stack:PythonTypeScriptJavaScriptGoAWSAzureGCPDockerKubernetesTerraformAWS CDKCloudFormationLinux/UnixGitOpenTelemetryPrometheusGrafanaPostgreSQLMySQLSQL

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor