Senior Database Reliability Engineer - DBRE

Cognite
Bengaluru
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 6+ yearsSkills: ["Ownership","Problem-solving","Troubleshooting","Incident management","Automation mindset"]

Own reliability, scalability, automation, and operational excellence for core database infrastructure. You’ll orchestrate a 1000+ PostgreSQL fleet across Azure, AWS, and GCP, run and scale Elasticsearch clusters (Elastic Cloud and ECK), and operate Kafka streaming infrastructure in both self-managed and managed configurations. Partner with Software Engineering, SRE, Platform Engineering, and Product to reduce operational toil and improve automated incident detection and recovery.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cognite
Cognite
1 month ago

Senior Database Reliability Engineer - DBRE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 42 minutes agoStatus: Live

Job Summary

Own reliability, scalability, automation, and operational excellence for core database infrastructure. You’ll orchestrate a 1000+ PostgreSQL fleet across Azure, AWS, and GCP, run and scale Elasticsearch clusters (Elastic Cloud and ECK), and operate Kafka streaming infrastructure in both self-managed and managed configurations. Partner with Software Engineering, SRE, Platform Engineering, and Product to reduce operational toil and improve automated incident detection and recovery.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Own reliability, scalability, automation, and operational excellence for core database infrastructure across PostgreSQL, Elasticsearch, and Kafka.
  • •Standardize and automate lifecycle management of a 1000+ PostgreSQL instance fleet across Azure, AWS, and GCP, including provisioning, upgrades, backups, and workflows.
  • •Design, operate, and scale Elasticsearch clusters across Elastic Cloud and ECK, covering capacity planning, shard management, upgrades, monitoring, and disaster recovery.
  • •Operate Kafka clusters across self-managed (Kubernetes-based, e.g., Strimzi/Kafka Operator) and managed (e.g., Confluent Cloud, MSK) configurations, including scaling, replication strategy, tuning, and upgrades.
  • •Partner with engineering teams to improve observability and build automation that reduces operational toil and enables early incident detection and resolution.

Key Requirements

  • •6+ years of experience in Database Reliability Engineering, Database Engineering, SRE, Platform Engineering, or a closely related role.
  • •Hands-on experience operating PostgreSQL at scale, preferably with cloud-managed environments.
  • •Strong experience with Elasticsearch, including cluster administration, performance tuning, scaling, shard management, and troubleshooting.
  • •Experience operating databases and stateful workloads on Kubernetes.
  • •Proficiency in scripting/programming with Python, Go, or a similar language, plus strong Infrastructure as Code experience with Terraform or similar tools.
Experience:6+ yearsDatabase reliability engineeringSREPlatform engineeringMulti-cloudDistributed systems
Skills:OwnershipProblem-solvingTroubleshootingIncident managementAutomation mindset
Tech Stack:PostgreSQLElasticsearchKafkaPythonGoTerraformKubernetesAzureAWSGCPElastic CloudECKStrimziKafka OperatorConfluent CloudMSKSchema RegistryInfrastructure as Code (IaC)

Company Brief

Cognite
Provides industrial data software that helps companies connect, contextualize, and use operational data for analytics, AI, and digital transformation. Its platform is used by heavy-industry customers to improve asset performance and operations.
Industry: Data Infrastructure
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Oslo, Norway
Founded: 2016
WebsiteLinkedIn