Staff Site Reliability Engineer

MongoDB
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical leadership","Mentoring","Collaboration","Automation mindset","Customer-focused"]

Provide technical leadership for the reliability architecture of a new platform enabling customers to build AI applications with MongoDB. Own operational foundations at scale across regions and multi-cloud providers, including capacity planning, incident response, and SLO discipline. Lead and mentor the SRE team responsible for Kubernetes fleet operations, networking, observability/alerting, and tenant isolation, while participating in a 24/7 on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
MongoDB
MongoDB
1 month ago

Staff Site Reliability Engineer

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Provide technical leadership for the reliability architecture of a new platform enabling customers to build AI applications with MongoDB. Own operational foundations at scale across regions and multi-cloud providers, including capacity planning, incident response, and SLO discipline. Lead and mentor the SRE team responsible for Kubernetes fleet operations, networking, observability/alerting, and tenant isolation, while participating in a 24/7 on-call rotation.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own the platform’s reliability architecture across regions and cloud providers.
  • •Set technical direction for platform operations including capacity planning, multi-cloud expansion, incident response, and SLO discipline.
  • •Collaborate with platform teams to provide guidance on operability, capacity, and best practices.
  • •Set operational standards for the team, including on-call quality, incident response, and SLO discipline.
  • •Mentor and technically develop the SRE team while participating in a 24/7 on-call rotation to resolve platform infrastructure issues.

Pay and Benefits

Perks:Parental LeaveFertility Assistance

Key Requirements

  • •10+ years operating distributed systems with deep Kubernetes expertise, including designing or evolving multi-cluster platforms.
  • •Proficiency in Python, Go, or a similar programming language.
  • •Ability to reason about workload isolation at the systems level (containers vs. virtual machines) for running untrusted code.
  • •Customer-focused mindset and preference for automation over manual processes.
  • •Familiarity with infrastructure primitives of at least one of AWS, GCP, or Azure, and ability to compare differences between them.
Experience:Distributed systemsCloudMulti-cloud
Skills:Technical leadershipMentoringCollaborationAutomation mindsetCustomer-focused
Languages:English
Tech Stack:PythonGoKubernetesAWSGCPAzureContainersVirtual machinesObservabilityAlertingSLOMulti-cluster platforms

Company Brief

MongoDB
Develops MongoDB, a leading general-purpose, document-based database platform that enables developers and enterprises to build scalable, high-performance applications with flexible data models and cloud-native capabilities.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2007
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn