Site Reliability Engineer

Dimension Data
Bucharest
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Collaborative teamwork","Proactive attitude","Self confidence","Attention to detail","Communication skills"]

Implement and operate an Online Banking platform on Google Cloud using the “you built it, you run it” model. Own SLOs/SLIs, improve resilience with reliability patterns (autoscaling, circuit breakers, rate limiting, retries), and drive incident, problem, and service request management to reduce toil and MTTR. Collaborate with analysts and architects, lead the SRE community, and advance observability and chaos engineering practices (OpenTelemetry, distributed tracing).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Dimension Data
Dimension Data
2 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 days agoStatus: Live

Job Summary

Implement and operate an Online Banking platform on Google Cloud using the “you built it, you run it” model. Own SLOs/SLIs, improve resilience with reliability patterns (autoscaling, circuit breakers, rate limiting, retries), and drive incident, problem, and service request management to reduce toil and MTTR. Collaborate with analysts and architects, lead the SRE community, and advance observability and chaos engineering practices (OpenTelemetry, distributed tracing).
Location: Bucharest
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Define SLOs/SLIs and enable an end-to-end customer satisfaction view to improve system performance and availability.
  • •Collaborate with business functional analysts and solution architects to improve resilience in solution design early on.
  • •Guide squads on reliability improvement prioritization and deliver improvements as part of sprints.
  • •Implement reliability and resilience patterns (auto-scaling, circuit breakers, bulkheads, rate limiting, retry mechanisms) and actively manage incident/problem/service request handling to reduce toil and MTTR.
  • •Advance SRE best practices (distributed tracing, OpenTelemetry, chaos engineering) and lead the SRE chapter population by establishing events/procedures and mentoring engineers.

Pay and Benefits

Perks:Health InsuranceRemote WorkLearning Budget

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or related field.
  • •Minimum 5 years of proven experience as a Reliability Engineer (or similar).
  • •Hands-on experience with Google Cloud and cloud-hosted applications, plus Docker/Kubernetes (including GKE) and Terraform (or similar).
  • •Experience building resilient software in Python/Java with modern CI/CD pipelines (e.g., GitHub/GitHub Actions/Bitbucket/Helm).
  • •Strong observability and self-healing experience (e.g., New Relic, Splunk, Google Cloud Operations, Lightstep, Ansible) and knowledge of security standards and microservices/APIs (e.g., TLS, OAuth2, KMS, Vault, Apigee/WSO2).
Experience:5+ years
Education:Bachelor's
Skills:Collaborative teamworkProactive attitudeSelf confidenceAttention to detailCommunication skills
Languages:English
Tech Stack:Google Cloud PlatformGKEDockerKubernetesTerraformPythonJavaCI/CDGitHubGitHub ActionsBitbucketHelmNew RelicSplunkGoogle Cloud OperationsLightstepAnsibleTLSOAuth2KMS

Company Brief

Dimension Data
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn