Senior Site Reliability Engineer, Fleet Management

MongoDB
Austin, Boston, Los Angeles, New York, Raleigh, San Francisco, Dublin
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsSkills: ["Communication","Ownership","Problem-solving","Collaboration","Blameless postmortems"]

Lead and maintain a scalable, secure Kubernetes-based runtime environment for MongoDB’s Fleet Management, collaborating with engineers to solve domain-specific problems, participating in 24/7 on-call rotations, and driving systemic fixes while migrating infrastructure from Terraform IaC to an Operator-driven lifecycle model.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
MongoDB
MongoDB
4 months ago

Senior Site Reliability Engineer, Fleet Management

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead and maintain a scalable, secure Kubernetes-based runtime environment for MongoDB’s Fleet Management, collaborating with engineers to solve domain-specific problems, participating in 24/7 on-call rotations, and driving systemic fixes while migrating infrastructure from Terraform IaC to an Operator-driven lifecycle model.
Location: Austin, Boston, Los Angeles, New York, Raleigh, San Francisco, Dublin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB.
  • •Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems.
  • •Participate in a 24/7 on-call rotation to resolve critical issues.
  • •Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice.
  • •Drive the migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model.

Pay and Benefits

Equity and Bonus:Equity
Perks:Equity401kPar​ental LeaveHealth Insurance

Key Requirements

  • •Have 6+ years of experience in software development and operating distributed systems.
  • •Proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (unit, integration, and E2E tests).
  • •Deep experience using and extending containerization technologies, preferably Kubernetes.
  • •Solid understanding of Linux operating system internals and networking concepts (filesystems, TCP/IP, DNS, TLS).
  • •Strong operational ownership, including debugging complex production issues and driving them to resolution.
Experience:6+ yearsKubernetesDistributed systemsGoPythonLinux
Skills:CommunicationOwnershipProblem-solvingCollaborationBlameless postmortems
Languages:English
Tech Stack:GoPythonKubernetesTerraformCrossplaneAWSGCPAzureHelmKustomizeGatekeeperKyvernoCRDsOperatorsLinuxCoreDNSCert-managerDNSTLS

Company Brief

MongoDB
Develops MongoDB, a leading general-purpose, document-based database platform that enables developers and enterprises to build scalable, high-performance applications with flexible data models and cloud-native capabilities.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2007
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn