Site Reliability Engineer

ArangoDB
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Troubleshooting","Problem-solving","Communication","Collaboration","Independence"]

Build and run ArangoDB’s cloud-native infrastructure that powers distributed database systems. As an SRE, you’ll design and maintain scalable AWS and Google Cloud environments and Kubernetes-based services, improve observability with monitoring/logging/alerting, and optimize CI/CD pipelines. You’ll automate repetitive operations using clean Go code, troubleshoot complex production incidents, and help drive high availability, fault tolerance, and disaster recovery through continuous improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ArangoDB
ArangoDB
1 day ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and run ArangoDB’s cloud-native infrastructure that powers distributed database systems. As an SRE, you’ll design and maintain scalable AWS and Google Cloud environments and Kubernetes-based services, improve observability with monitoring/logging/alerting, and optimize CI/CD pipelines. You’ll automate repetitive operations using clean Go code, troubleshoot complex production incidents, and help drive high availability, fault tolerance, and disaster recovery through continuous improvements.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design, implement, and maintain cloud infrastructure on AWS and Google Cloud for distributed database systems.
  • •Ensure scalability, performance, and reliability of Kubernetes-based distributed database services.
  • •Collaborate with developers to write production-grade Go code to automate infrastructure management and improve operations.
  • •Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems for production.
  • •Participate in on-call rotations and troubleshoot incidents; implement monitoring, logging, and alerting for system health and performance.

Key Requirements

  • •Proven experience as an SRE or DevOps Engineer in a cloud-native environment.
  • •Proficiency with Kubernetes for large-scale, distributed systems.
  • •Experience with cloud providers including AWS and Google Cloud (GCP).
  • •Strong Linux and networking fundamentals, with troubleshooting skills for complex infrastructure issues.
  • •Knowledge of CI/CD practices and observability tooling (e.g., Prometheus, Grafana, ELK), plus experience with automation/code in Golang or willingness to learn.
Experience:Cloud-nativeKubernetesDevOpsDistributed systems
Skills:TroubleshootingProblem-solvingCommunicationCollaborationIndependence
Tech Stack:KubernetesAWSGoogle CloudGCPGolangGoPythonDockerLinuxCI/CDJenkinsCircleCIPrometheusGrafanaELK stackGitBashTerraformGitOps

Company Brief

ArangoDB
ArangoDB builds a multi-model database combining graph, document, and key-value models with a unified query language and managed cloud (Oasis), enabling scalable graph analytics, search, and ML use cases for enterprises.
Industry: Data Infrastructure
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Private Equity Backed
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor