Senior Site Reliability Engineer

Gradle
North America
Workplace: RemoteFull timeUSD 150,000 - 190,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Self-direction","Written communication","Incident management","Automation mindset","Clear communication"]

Help found a new SRE team by operating and improving Develocity’s production services. You’ll troubleshoot incidents across the stack, drive automation for deployment, upgrades, monitoring, and recovery, and build observability with logging, metrics, tracing, and alerting. Working with AWS Kubernetes and the internal Cloud Application Platform, you’ll embed reliability into how software is shipped and communicate with customers during incidents and maintenance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Gradle
Gradle
1 month ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Help found a new SRE team by operating and improving Develocity’s production services. You’ll troubleshoot incidents across the stack, drive automation for deployment, upgrades, monitoring, and recovery, and build observability with logging, metrics, tracing, and alerting. Working with AWS Kubernetes and the internal Cloud Application Platform, you’ll embed reliability into how software is shipped and communicate with customers during incidents and maintenance.
Location: North America
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate and maintain Develocity instances and supporting services.
  • •Participate in follow-the-sun on-call and troubleshoot issues across the stack.
  • •Drive automation across deployment, upgrades, monitoring, self-healing, and recovery.
  • •Build and maintain observability (logging, metrics, tracing, alerting) for managed services.
  • •Own disaster recovery, backups, business continuity, and run incident response and retrospectives.

Pay and Benefits

Salary: USD 150,000 - 190,000 annually
Equity and Bonus:Equity
Perks:Remote WorkEquity Grants

Key Requirements

  • •5+ years in SRE, DevOps, or an equivalent role operating production services at scale.
  • •Strong Kubernetes experience in production environments.
  • •Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
  • •Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
  • •Track record of incident management, including 24/7 on-call rotations and knowledge of SRE best practices (SLAs, SLOs).
Experience:5+ yearsSaaSProduction operationsObservabilitySREDevOps
Skills:Self-directionWritten communicationIncident managementAutomation mindsetClear communication
Languages:English
Tech Stack:KubernetesAWSEKSRDSS3EC2PrometheusGrafanaTerraformPythonBashCloud Application PlatformObservabilityLoggingMetricsTracingAlertingInfrastructure as Code

Company Brief

Gradle
Develops Gradle, a build automation and dependency management tool and Gradle Enterprise for optimizing build and test performance. Serves software engineering teams to accelerate CI/CD, improve developer productivity, and scale build infrastructure.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: San Francisco, United States
Founded: 2007
WebsiteLinkedIn