Senior Site Reliability Engineer, Compute

Roblox
San Mateo
Workplace: OnsiteFull timeUSD 243,290 - 295,250 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsEducation: bachelorsSkills: ["Curiosity","Collaboration","Problem-solving","Planning","Data-driven thinking"]

Own and run the Roblox Compute infrastructure cell platform, including service discovery, secrets management, and related software layers. Build and standardize a “golden path” for private-cloud cluster operations, production guardrails, and load-testing-based release readiness. Create performance monitoring and observability to detect capacity issues and platform degradations, while analyzing designs for production readiness. Lead reliability best practices across the Infra Compute group through reviews and operational guidance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Roblox
Roblox
1 day ago

Senior Site Reliability Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live
Reposted: similar role first listed 9 months ago

Job Summary

Own and run the Roblox Compute infrastructure cell platform, including service discovery, secrets management, and related software layers. Build and standardize a “golden path” for private-cloud cluster operations, production guardrails, and load-testing-based release readiness. Create performance monitoring and observability to detect capacity issues and platform degradations, while analyzing designs for production readiness. Lead reliability best practices across the Infra Compute group through reviews and operational guidance.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and develop systems and libraries that improve fault-tolerance and resilience, automate cluster management/lifecycle tasks, and ensure observability.
  • •Institute reliability best practices across the Infra Compute group and drive common reliability initiatives through technical reviews and operational guidance.
  • •Build, automate, and standardize process automation to create a “golden path” of tooling and platform support for Roblox’s ecosystem.
  • •Create production guardrails by evaluating release-candidate capacity with load-testing tooling before deploying to production.
  • •Develop performance monitoring and observability to understand capacity issues and platform degradations, including alerting and canarying, and analyze production readiness.

Pay and Benefits

Salary: USD 243,290 - 295,250 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s degree (or equivalent professional experience) in Computer Science or a related engineering field, with at least 6 years as an SRE or Software Engineer.
  • •7+ years of professional experience with high-level programming languages such as Go, Java, or C#.
  • •Experience with Kubernetes or similar orchestration systems.
  • •Strong desired experience with Nomad, Vault, and Consul.
  • •Good habits for building software and tools and driving adoption, with a focus on deeply reliable systems.
Experience:6+ yearsInfrastructureCloud infrastructureKubernetes
Education:Bachelor's in Computer Science or related engineering field
Skills:CuriosityCollaborationProblem-solvingPlanningData-driven thinking
Languages:English
Tech Stack:GoJavaC#KubernetesNomadVaultConsul

Eligibility

Visa:H-1B

Company Brief

Roblox
Operates an online platform for user-generated games and virtual experiences, enabling creators to build, publish, and monetize interactive 3D experiences for a global community of players.
Industry: Gaming
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Mateo, United States
Founded: 2004
WebsiteLinkedIn