Senior Site Reliability Engineer, Compute

Roblox
San Mateo
Workplace: OnsiteFull timeUSD 243,290 - 295,250 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsEducation: bachelorsSkills: ["Collaboration","Problem-solving","Planning","Curiosity","Data-driven thinking"]

Own and operate the Infrastructure Compute cell infrastructure system and related layers like service discovery and secrets management. Build Roblox’s private cloud, productionize Kubernetes-based infrastructure, and drive reliability best practices across the Compute team. Create fault-tolerant, observable tooling and libraries, automate cluster lifecycle processes, and implement production guardrails using load testing, monitoring, and canarying to improve capacity and service reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Roblox
Roblox
2 months ago

Senior Site Reliability Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live
Reposted: similar role first listed 9 months ago

Job Summary

Own and operate the Infrastructure Compute cell infrastructure system and related layers like service discovery and secrets management. Build Roblox’s private cloud, productionize Kubernetes-based infrastructure, and drive reliability best practices across the Compute team. Create fault-tolerant, observable tooling and libraries, automate cluster lifecycle processes, and implement production guardrails using load testing, monitoring, and canarying to improve capacity and service reliability.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and develop systems and libraries that promote fault-tolerance and resilience, automate cluster management/lifecycle, and ensure observability.
  • •Promote and institute reliability best practices across the Infra Compute group, including technical reviews and operational guidance.
  • •Build, automate, and standardize tooling and processes to create a “golden path” for platform support across Roblox’s ecosystem.
  • •Create production guardrails by evaluating release candidate capacity with load testing tooling before deployment.
  • •Create performance monitoring and observability to identify capacity issues and platform degradations, including canarying services with alerting and production change monitoring.

Pay and Benefits

Salary: USD 243,290 - 295,250 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor degree (or equivalent professional experience) in Computer Science or related engineering field with at least 6 years as an SRE or Software Engineer.
  • •Fluency with high-level programming languages such as Go, Java, and C#.
  • •Experience with Kubernetes or similar orchestration systems; Nomad, Vault, and Consul are strongly desired.
  • •Strong habits for building software/tools and getting them adopted across teams, with a focus on deeply reliable code.
  • •Experience in building systems and analyzing designs for production readiness.
Experience:6+ years
Education:Bachelor's in Computer Science
Skills:CollaborationProblem-solvingPlanningCuriosityData-driven thinking
Tech Stack:GoJavaC#KubernetesNomadVaultConsulLoad testingObservabilityCanarying

Eligibility

Visa:H-1B

Company Brief

Roblox
Operates an online platform for user-generated games and virtual experiences, enabling creators to build, publish, and monetize interactive 3D experiences for a global community of players.
Industry: Gaming
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Mateo, United States
Founded: 2004
WebsiteLinkedIn