Senior Site Reliability Engineer

Duolingo
Pittsburgh, New York
Workplace: OnsiteFull timeUSD 182,800 - 247,300 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Collaboration","Problem-solving","Operational excellence","Incident response","System design consulting"]

Own the reliability of Duolingo’s distributed systems by partnering with product and platform engineering teams. Identify sources of instability, support production operations, and provide system design guidance. Drive incident response and postmortem practices, improve reliability and scalability through continuous change, and reduce operational toil with automation. Become an authority on Duolingo services while collaborating across teams to launch new features.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Duolingo
Duolingo
3 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own the reliability of Duolingo’s distributed systems by partnering with product and platform engineering teams. Identify sources of instability, support production operations, and provide system design guidance. Drive incident response and postmortem practices, improve reliability and scalability through continuous change, and reduce operational toil with automation. Become an authority on Duolingo services while collaborating across teams to launch new features.
Location: Pittsburgh, New York
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence.
  • •Support core infrastructure by understanding, diagnosing, and debugging production systems.
  • •Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis.
  • •Maintain and document incident response and postmortem practices.
  • •Implement changes to improve reliability, scalability, and velocity, including tooling and automation to reduce toil.

Pay and Benefits

Salary: USD 182,800 - 247,300 annually

Key Requirements

  • •5+ years of experience in site reliability engineering/DevOps for a product with millions of users.
  • •Experience identifying and solving issues in large-scale distributed systems.
  • •Experience with Java, Kotlin, Python, or Go.
  • •Knowledge of containerization and container orchestration tools such as Docker, Mesos, Kubernetes, or Nomad.
  • •Familiarity with reliability troubleshooting in Dynamo and/or MySQL and/or PostgreSQL databases.
Experience:5+ years
Skills:CollaborationProblem-solvingOperational excellenceIncident responseSystem design consulting
Languages:En
Tech Stack:JavaKotlinPythonGoDockerMesosKubernetesNomadDynamoMySQLPostgreSQL

Company Brief

Duolingo
Duolingo develops a gamified language-learning platform offering courses, practice exercises, and assessment tools for learners worldwide via web and mobile apps, using adaptive algorithms and motivational features to drive engagement and retention.
Industry: EdTech
Company Size: Enterprise (1,001+ employees)
Revenue: USD 500M to 1B
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Pittsburgh, United States
Founded: 2011
Glassdoor
Glassdoor: 4.2
WebsiteLinkedInGlassdoor