Staff Site Reliability Engineer

Gradle
Europe
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Mentoring","Self-direction"]

Help build a new SRE team and lead reliability for Develocity’s production services, including paying customer, open-source, and public-facing workloads. Define SRE standards and operating models (on-call, incident response, postmortems, SLOs), drive automation and observability, and troubleshoot across the stack. You’ll work on a cloud application platform running Kubernetes on AWS and partner closely with engineering and cloud platform teams to balance reliability with delivery excellence.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Gradle
Gradle
1 month ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Help build a new SRE team and lead reliability for Develocity’s production services, including paying customer, open-source, and public-facing workloads. Define SRE standards and operating models (on-call, incident response, postmortems, SLOs), drive automation and observability, and troubleshoot across the stack. You’ll work on a cloud application platform running Kubernetes on AWS and partner closely with engineering and cloud platform teams to balance reliability with delivery excellence.
Location: Europe
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and maintain Develocity instances and supporting services in production
  • •Define and evolve SRE standards, practices, and operating models (on-call, incident response, postmortems, SLOs)
  • •Lead incident response and blameless retrospectives to deliver measurable reliability improvements
  • •Drive automation across deployment, upgrades, monitoring, self-healing, recovery, and operational workflows
  • •Build and maintain observability and own disaster recovery, backups, and business continuity

Pay and Benefits

Equity and Bonus:Equity
Perks:Remote WorkEquity

Key Requirements

  • •7+ years in SRE, DevOps, or an equivalent role operating production services at scale
  • •Experience leading reliability initiatives across multiple teams or services
  • •Proficiency designing and operating systems with SLOs and error budgets while balancing reliability, velocity, and cost
  • •Strong Kubernetes production experience
  • •Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2), plus observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform)
Experience:SREDevOpsSaaSCloud infrastructureKubernetesObservability
Skills:CommunicationMentoringSelf-direction
Languages:English
Tech Stack:KubernetesAWSEKSRDSS3EC2PrometheusGrafanaTerraformPythonBashLoggingMetricsTracingAlertingInfrastructure as Code

Company Brief

Gradle
Develops Gradle, a build automation and dependency management tool and Gradle Enterprise for optimizing build and test performance. Serves software engineering teams to accelerate CI/CD, improve developer productivity, and scale build infrastructure.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: San Francisco, United States
Founded: 2007
WebsiteLinkedIn