Senior Site Reliability Engineer, Database Infrastructure

Zello
Austin
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsEducation: bachelorsSkills: ["Python","Go","Bash","Docker","Kubernetes","Prometheus","OpenTelemetry","Git","SQL","NoSQL","MySQL","MongoDB","ScyllaDB","Elasticsearch","Redis","Loki","Tempo","AI tooling"]

Senior Site Reliability Engineer who owns the data tier reliability (MySQL, MongoDB, ScyllaDB, Elasticsearch, Redis) and contributes to monitoring, on-call, and cloud modernization, with a focus on performance, observability, incident response, and automation across production databases in a hybrid workplace.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zello
Zello
3 months ago

Senior Site Reliability Engineer, Database Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Senior Site Reliability Engineer who owns the data tier reliability (MySQL, MongoDB, ScyllaDB, Elasticsearch, Redis) and contributes to monitoring, on-call, and cloud modernization, with a focus on performance, observability, incident response, and automation across production databases in a hybrid workplace.
Location: Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design, deploy, and operate highly available MySQL and MongoDB clusters across cloud environments, including replication, sharding, backups, point-in-time recovery, upgrades, and disaster recovery
  • •Tune query performance, schema, and index strategy in collaboration with application engineers and push fixes upstream when applicable
  • •Extend observability stack (Prometheus, Loki, Tempo) to ensure data tier instrumentation matches the application tier
  • •Participate in the Platform on-call rotation, lead incident response for data-tier issues, and write postmortems to drive durable change
  • •Improve disaster recovery, security posture, and compliance for the database footprint (encryption, access control, audit logging, backup integrity)

Pay and Benefits

Perks:EquityFree SnacksSabbaticalTime Off

Key Requirements

  • •7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, with at least 3 years owning production databases
  • •BSc in Computer Science or equivalent practical experience
  • •Experience operating highly available MySQL and MongoDB in production at scale, including replication, sharding, backups, point-in-time recovery, and failover drills
  • •Proficiency in automation using Python, Go, or Bash for tooling and operators
  • •Experience with containerized workloads (Docker/Kubernetes) and major cloud providers (GCP preferred; AWS/Azure acceptable) and strong observability skills (Prometheus/OpenTelemetry, dashboards)
Experience:7+ years
Education:Bachelor's
Skills:PythonGoBashDockerKubernetesPrometheusOpenTelemetryGitSQLNoSQLMySQLMongoDBScyllaDBElasticsearchRedisLokiTempoAI tooling
Tech Stack:PythonGoBashDockerKubernetesPrometheusLokiTempoMySQLMongoDBScyllaDBElasticsearchRedisOpenTelemetryGitSQLCloudGCPAWSAzure

Company Brief

Zello
Provides a push-to-talk (PTT) voice messaging platform and real-time communication apps for frontline teams and communities, enabling instant voice communication over cellular and Wi-Fi networks.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Established Company
Headquarters: Austin, United States
Founded: 2007
WebsiteLinkedIn