Site Reliability Engineer

Blaxel
San Francisco
Workplace: OnsiteFull timeUSD 175,000 - 250,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Problem-solving","Communication","Analytical","Automation","Teamwork"]

We are seeking a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of an AI infrastructure platform. You’ll own end-to-end reliability, from observability and incident response to automation and security, architecting systems for ultra-low latency workloads and billions of agent requests, while collaborating with founders, infra, and dev teams in a fast-growing infra environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Blaxel
Blaxel
4 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

We are seeking a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of an AI infrastructure platform. You’ll own end-to-end reliability, from observability and incident response to automation and security, architecting systems for ultra-low latency workloads and billions of agent requests, while collaborating with founders, infra, and dev teams in a fast-growing infra environment.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Architect, operate, and continuously improve the core infrastructure powering our 25ms cold-start compute engine.
  • •Build and evolve our observability stack (metrics, traces, logs), ensuring we detect issues before users do.
  • •Define, monitor, and drive SLOs/SLIs across key system surfaces to maintain world-class reliability.
  • •Lead incident response with rigor: root cause analysis, post-mortems, and driving systemic fixes.
  • •Design and implement self-healing, automated operational systems to eliminate toil and scale ops.

Pay and Benefits

Salary: USD 175,000 - 250,000 annually

Key Requirements

  • •3+ years in SRE, DevOps, or infrastructure engineering roles
  • •Strong proficiency in Go, Rust, or Python
  • •Hands-on experience with AWS or GCP
  • •Solid knowledge of Linux systems, networking, and distributed systems
  • •Experience with Kubernetes or similar orchestrators
Experience:3+ yearsAICloud computingSaaS
Skills:Problem-solvingCommunicationAnalyticalAutomationTeamwork
Tech Stack:GoRustPythonAWSGCPLinuxKubernetesPrometheusGrafanaELKDatadogGitHub ActionsGitLab CIJenkinsTerraformPulumi

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Blaxel
Provides AI-powered tools for creating, customizing, and managing visual content to help teams accelerate design workflows and produce brand-aligned imagery at scale.
Industry: SaaS
Website