Senior Staff Site Reliability Engineer

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Technical leadership","Mentorship","Building standards","Cross-functional influence","Performance improvement"]

Design and deliver a scalable enterprise AI runtime platform, defining architecture and technical roadmaps for deploying, operating, and scaling AI applications and inference services across cloud and on-prem environments. Build Kubernetes-based systems, control-plane services, APIs, operators, and automation for provisioning and recovery. Enhance GPU scheduling, autoscaling, and low-latency inference performance while improving availability, observability, and application lifecycle management.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live
Reposted: similar role first listed 1 week ago

Job Summary

Design and deliver a scalable enterprise AI runtime platform, defining architecture and technical roadmaps for deploying, operating, and scaling AI applications and inference services across cloud and on-prem environments. Build Kubernetes-based systems, control-plane services, APIs, operators, and automation for provisioning and recovery. Enhance GPU scheduling, autoscaling, and low-latency inference performance while improving availability, observability, and application lifecycle management.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define the architecture and technical roadmap for a scalable enterprise AI runtime platform.
  • •Design Kubernetes-based systems for deploying, running, and scaling AI applications, inference services, and databases.
  • •Build control-plane services, APIs, operators, and automation for workload provisioning, configuration, upgrades, and recovery.
  • •Develop runtime capabilities for GPU scheduling, autoscaling, load balancing, and rate limiting to improve inference performance.
  • •Improve observability and application lifecycle-management patterns across cloud and on-prem platforms, mentoring engineers and setting long-term standards.

Key Requirements

  • •BS, MS, or PhD in Computer Science/Engineering or related field, or equivalent experience.
  • •8+ years building distributed systems, cloud infrastructure, database platforms, or large-scale backend services.
  • •Strong programming skills in Python, Go, C++, or Java for production-grade systems.
  • •Experience designing scalable, highly available Kubernetes-based platforms and leading technical strategy across teams.
  • •Hands-on experience with control planes, platform APIs, Kubernetes operators, or workload lifecycle-management systems.
Experience:8+ yearsDistributed systemsCloud infrastructureKubernetesDatabase platformsAI inference
Education:Bachelor's in Computer Science, Engineering
Skills:Technical leadershipMentorshipBuilding standardsCross-functional influencePerformance improvement
Tech Stack:PythonGoC++JavaKubernetesAPIsGitOpsCI/CDObservabilityGPU schedulingAutoscalingLoad balancingRate limitingRelational databasesVector databasesBackupsFailoverProfilingDebugging

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor