Site Reliability Engineer

ArangoDB
Spain, Madrid
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Troubleshooting","Problem-solving","Communication","Collaboration","Self-organization"]

Build and run cloud-native infrastructure for Arango’s distributed database systems. You’ll design and maintain reliable Kubernetes workloads on AWS and Google Cloud, improve observability, and automate CI/CD and deployment workflows. Partner with developers to write production-grade automation code in Golang, troubleshoot complex cross-stack issues, and contribute to disaster recovery, high availability, and fault tolerance while participating in on-call rotations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ArangoDB
ArangoDB
3 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live
Reposted: similar role first listed 2 days ago

Job Summary

Build and run cloud-native infrastructure for Arango’s distributed database systems. You’ll design and maintain reliable Kubernetes workloads on AWS and Google Cloud, improve observability, and automate CI/CD and deployment workflows. Partner with developers to write production-grade automation code in Golang, troubleshoot complex cross-stack issues, and contribute to disaster recovery, high availability, and fault tolerance while participating in on-call rotations.
Location: Spain, Madrid
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design, implement, and maintain cloud infrastructure on AWS and Google Cloud.
  • •Maintain scalability, performance, and reliability of Kubernetes-based distributed database systems.
  • •Collaborate with developers to write production-grade Golang code to automate infrastructure management and improve operations.
  • •Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems for production.
  • •Develop disaster recovery, high availability, and fault-tolerance strategies and participate in on-call rotations to resolve incidents.

Key Requirements

  • •Proven experience as an SRE or DevOps engineer in a cloud-native environment.
  • •Strong Kubernetes experience managing large-scale, distributed systems.
  • •Experience with AWS and Google Cloud (GCP), plus solid networking/security troubleshooting knowledge.
  • •Understanding of Linux internals and containerization (e.g., Docker).
  • •Knowledge of CI/CD and observability/monitoring tools (e.g., Prometheus, Grafana, ELK) and the ability to troubleshoot complex infrastructure issues.
Experience:Cloud-nativeKubernetesDistributed systemsDistributed databasesObservability
Skills:TroubleshootingProblem-solvingCommunicationCollaborationSelf-organization
Tech Stack:KubernetesAWSGoogle CloudGCPGolangPythonLinuxDockerJenkinsCircleCIPrometheusGrafanaELK stackGitBashTerraformGitOps

Company Brief

ArangoDB
ArangoDB builds a multi-model database combining graph, document, and key-value models with a unified query language and managed cloud (Oasis), enabling scalable graph analytics, search, and ML use cases for enterprises.
Industry: Data Infrastructure
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Private Equity Backed
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor