Staff Site Reliability Engineer - Site Experience

Reddit
Dublin
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Go","Python","Kubernetes","Prometheus","Grafana","Opentelemetry","Envoy","Kafka","Redis","Cassandra","Logging","Monitoring"]

Lead reliability engineering for user-facing systems at internet scale, partnering with product and infrastructure teams to improve availability, latency, and operational excellence. Architect for scale, reduce operational risk, drive automation, manage incidents, and mentor teams to raise the bar on reliability across Reddit’s real-time services.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reddit
Reddit
3 months ago

Staff Site Reliability Engineer - Site Experience

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Lead reliability engineering for user-facing systems at internet scale, partnering with product and infrastructure teams to improve availability, latency, and operational excellence. Architect for scale, reduce operational risk, drive automation, manage incidents, and mentor teams to raise the bar on reliability across Reddit’s real-time services.
Location: Dublin
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Lead Reliability Engineering for User Experience; Drive reliability, scalability, and operational excellence for critical user facing systems and services.
  • •Architect for Scale; Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load.
  • •Reduce Operational Risk; Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure.
  • •Drive Automation; Eliminate repetitive operational work through automation and tooling.
  • •Incident Management; Lead complex incident response efforts across engineering teams and drive blameless postmortems.

Pay and Benefits

Perks:Private MedicalDentalVisionRetirement SavingsPaid LeaveParental LeaveHealth Insurance

Key Requirements

  • •8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems.
  • •Strong collaboration and communication skills with the ability to influence technical direction across teams.
  • •Experience designing highly available systems with strong operational and reliability practices.
  • •Strong programming skills in languages such as Go, Python, or similar.
  • •Deep understanding of distributed systems, networking, Linux systems, cloud native architectures.
Experience:8+ years
Skills:GoPythonKubernetesPrometheusGrafanaOpentelemetryEnvoyKafkaRedisCassandraLoggingMonitoring
Languages:English
Tech Stack:GoPythonKubernetesContainersCloudPrometheusGrafanaOpenTelemetryEnvoyKafkaClickHouseCassandraRedis

Company Brief

Reddit
Operates Reddit, a large online community and discussion platform where users submit content, comment, and vote across topic-based communities (subreddits); monetizes via advertising, premium subscriptions, and awards.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2005
WebsiteLinkedIn