Staff Site Reliability Engineer

Filevine
United States
Workplace: RemoteFull timeUSD 235,000 - 275,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 12+ yearsSkills: ["Mentorship","Technical leadership","Communication","Strategic partnership","Problem-solving"]

Own the technical standard and roadmap for Observability & Alerting and Platform Infrastructure, ensuring reliability problems are solved permanently. Serve as the senior technical authority for production operations—setting SLIs/SLOs, incident response practices, capacity planning, and automation. Build self-service platform capabilities to reduce toil and improve engineering safety. Lead through complex incidents, turn learning into lasting improvements, and mentor engineers across the organization.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Filevine
Filevine
1 month ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Own the technical standard and roadmap for Observability & Alerting and Platform Infrastructure, ensuring reliability problems are solved permanently. Serve as the senior technical authority for production operations—setting SLIs/SLOs, incident response practices, capacity planning, and automation. Build self-service platform capabilities to reduce toil and improve engineering safety. Lead through complex incidents, turn learning into lasting improvements, and mentor engineers across the organization.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and execute the technical strategy for Observability & Alerting and Platform Infrastructure and operational excellence.
  • •Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • •Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation across the service lifecycle.
  • •Lead complex production incidents and convert post-incident learning into permanent engineering improvements.
  • •Mentor engineers and serve as a trusted technical authority for long-term reliability and platform direction.

Pay and Benefits

Salary: USD 235,000 - 275,000 annually
Perks:Health InsuranceDentalVisionParental LeaveDisability

Key Requirements

  • •12+ years in software engineering, infrastructure, platform engineering, or SRE, including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives for distributed production systems.
  • •Expert-level depth in observability and platform infrastructure, including incident response, capacity planning, automation, and reliability engineering.
  • •Advanced experience with a major container orchestration platform, preferably Kubernetes, and an observability platform such as New Relic or Datadog.
  • •Strong software engineering ability in Python, Go, Bash, or another general-purpose language, with experience building production tooling and automation.
  • •Proven ability to mentor engineers and clearly communicate technical risk to engineering, product, and executive audiences.
Experience:12+ yearsSREDistributed systemsCloud infrastructureObservabilityAIOpsContainer orchestration
Skills:MentorshipTechnical leadershipCommunicationStrategic partnershipProblem-solving
Tech Stack:PythonGoBashInfrastructure as CodeKubernetesNew RelicDatadogObservabilityAIOps

Company Brief

Filevine
Provides cloud-based legal case management and practice management software for law firms and legal departments, offering document management, communication, reporting, and workflow automation to streamline case handling and client collaboration.
Industry: LegalTech
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Salt Lake City, United States
Founded: 2014
WebsiteLinkedIn