Senior Site Reliability Engineer, Platform Infrastructure (Foundations)

Anyscale
San Francisco, Palo Alto
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsEducation: bachelorsSkills: ["Go","Python","Kubernetes","Prometheus","Grafana","Linux"]

Senior SRE to design, build, and scale the control and data plane for Anyscale’s Ray-based platform. You’ll work on Kubernetes-based deployments, cloud-native infrastructure, and distributed AI workloads, with responsibilities spanning cluster orchestration, scheduling, reliability, observability, and on-call incident response, collaborating across open-source Ray tooling and Anyscale products.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anyscale
Anyscale
2 months ago

Senior Site Reliability Engineer, Platform Infrastructure (Foundations)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Senior SRE to design, build, and scale the control and data plane for Anyscale’s Ray-based platform. You’ll work on Kubernetes-based deployments, cloud-native infrastructure, and distributed AI workloads, with responsibilities spanning cluster orchestration, scheduling, reliability, observability, and on-call incident response, collaborating across open-source Ray tooling and Anyscale products.
Location: San Francisco, Palo Alto
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
  • •Optimize control plane components for large-scale, distributed AI/ML workloads
  • •Build intelligent scheduling and resource management systems for heterogeneous compute clusters
  • •Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
  • •Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues

Key Requirements

  • •Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
  • •3+ years of experience writing high-quality production code
  • •Hands-on experience in building and maintaining highly available, scalable, and performant distributed systems
  • •Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
  • •Proficiency in Go and Python
Experience:3+ yearsDistributed systemsCloud
Education:Bachelor's
Skills:GoPythonKubernetesPrometheusGrafanaLinux
Tech Stack:GoPythonKubernetesPrometheusGrafanaAWSAzureGCPLinuxContainers

Company Brief

Anyscale
Provides a managed platform and developer tooling built on the Ray open-source project to scale Python and AI applications across clouds, simplifying distributed training, inference and production deployment for engineering teams.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor