Senior Staff Site Reliability Engineer

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Technical leadership","Mentoring","Standards setting","Cross-functional collaboration","Problem-solving"]

Build the runtime foundation for enterprise AI platforms, providing technical leadership for systems that deploy, operate, and scale AI applications, inference services, and databases across cloud and on-prem environments. Own the architecture and roadmap for Kubernetes-based platforms, control-plane services, APIs, automation, and GPU-aware scheduling. Improve performance, availability, and developer experience while developing observability for GPU and inference workloads and leading initiatives across teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 days ago

Senior Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build the runtime foundation for enterprise AI platforms, providing technical leadership for systems that deploy, operate, and scale AI applications, inference services, and databases across cloud and on-prem environments. Own the architecture and roadmap for Kubernetes-based platforms, control-plane services, APIs, automation, and GPU-aware scheduling. Improve performance, availability, and developer experience while developing observability for GPU and inference workloads and leading initiatives across teams.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define the architecture and technical roadmap for a scalable enterprise AI runtime platform.
  • •Design Kubernetes-based systems for deploying, running, and scaling AI applications, inference services, and databases.
  • •Build control-plane services, APIs, operators, and automation for workload provisioning, configuration, upgrades, and recovery.
  • •Develop runtime capabilities for GPU scheduling, autoscaling, load balancing, and rate limiting.
  • •Lead technical initiatives across functions and mentor engineers while establishing platform standards for long-term evolution.

Key Requirements

  • •BS, MS, or PhD in Computer Science, Engineering, or a related field (or equivalent experience).
  • •8+ years of software engineering experience building distributed systems, cloud infrastructure, database platforms, or large-scale backend services.
  • •Strong programming skills in Python, Go, C++, or Java for production-grade systems.
  • •Proven experience designing scalable, highly available Kubernetes-based platforms and leading technical strategy across teams.
  • •Experience building control planes, platform APIs, Kubernetes operators, or workload lifecycle-management systems.
Experience:8+ yearsDistributed systemsCloud infrastructureKubernetesBackend servicesAI inferenceDatabasesCloud-native
Education:
Skills:Technical leadershipMentoringStandards settingCross-functional collaborationProblem-solving
Tech Stack:PythonGoC++JavaKubernetesGitOpsCI/CDObservabilityRelational databasesVector databases

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor