Site Reliability Engineer

HappyRobot
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Communication","Problem-solving","Collaboration","Attention to detail"]

Lead scaling of operational resilience by owning stability, observability, and debugging workflows. Tackle complex real-time incidents, design internal tooling to reduce incident load, and empower developers with reliable systems. Work in a fast-paced, high-trust environment to shift from reactive to proactive operations and improve uptime across critical services.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
HappyRobot
HappyRobot
8 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Lead scaling of operational resilience by owning stability, observability, and debugging workflows. Tackle complex real-time incidents, design internal tooling to reduce incident load, and empower developers with reliable systems. Work in a fast-paced, high-trust environment to shift from reactive to proactive operations and improve uptime across critical services.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own the stability, observability, and debugging workflows to keep systems running smoothly
  • •Untangle complex failures in real time and design tools that turn chaos into clarity
  • •Reduce incident load and improve developer focus and system uptime
  • •Build internal tooling for on-call reliability and proactive operations
  • •Collaborate with engineers to scale operational resilience as the company grows

Key Requirements

  • •3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)
  • •Strong problem-solving skills and ability to dive into unfamiliar backend codebases
  • •Comfort with Python and Go for reading code and writing small tools/utilities
  • •Familiarity with observability and monitoring tools (e.g., Datadog, Prometheus, Sentry)
  • •Clear, calm communication under pressure — especially during live incidents
Experience:3+ yearsAIEnterprise softwareSaaS
Skills:CommunicationProblem-solvingCollaborationAttention to detail
Languages:English
Tech Stack:PythonGoDatadogPrometheusSentry

Company Brief

HappyRobot
Builds and manages enterprise AI workers (conversational and document-processing agents) to automate operational workflows for supply chain and logistics companies, integrating with systems to execute tasks, collect real-time data, and drive operational insights.
Industry: Supply Chain Technology
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 250M to 500M
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2022
WebsiteLinkedInGlassdoor