Staff Site Reliability Engineer

BeyondTrust
Toronto, United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Mentorship","Technical leadership","Incident reviews","Documentation","Standards setting"]

Lead the evolution of the Password Safe platform by architecting highly available, resilient systems across cloud (AWS/Azure) and on-prem environments. Own reliability for shared services, CI/CD pipelines, and platform engineering initiatives, driving “Everything as Code” with GitOps automation. Build automation for CI/CD, infrastructure-as-code (Terraform/OpenTofu/Ansible), and observability using Grafana Cloud, Datadog, and OpenTelemetry, while mentoring engineers and raising SRE standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
BeyondTrust
BeyondTrust
6 days ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead the evolution of the Password Safe platform by architecting highly available, resilient systems across cloud (AWS/Azure) and on-prem environments. Own reliability for shared services, CI/CD pipelines, and platform engineering initiatives, driving “Everything as Code” with GitOps automation. Build automation for CI/CD, infrastructure-as-code (Terraform/OpenTofu/Ansible), and observability using Grafana Cloud, Datadog, and OpenTelemetry, while mentoring engineers and raising SRE standards.
Location: Toronto, United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, scale, and maintain highly available, secure, and resilient systems across cloud and on-prem environments.
  • •Standardize and optimize CI/CD pipelines for safe, repeatable, rapid deployments.
  • •Implement infrastructure-as-code and enforce an “Everything as Code” approach with version-controlled deployments via automated GitOps workflows.
  • •Architect telemetry (metrics, logs, traces) and define SLOs/SLIs for critical applications.
  • •Partner with engineering leadership to define long-term SRE strategy and mentor engineers through reliability practices and incident/design reviews.

Key Requirements

  • •7+ years in SRE/DevOps/Platform Engineering, including at least 2 years at Senior or Staff level.
  • •Demonstrated systems architecture and automation of infrastructure across cloud and on-prem environments.
  • •Experience containerizing and administering systems with Docker and Kubernetes, including networking and security primitives.
  • •Proficiency in release orchestration strategies such as Canary or Blue Green deployments.
  • •Strong software engineering ability in at least one systems language (Go, Java, or C#) to build automation, tooling, and internal APIs.
Experience:7+ yearsCybersecurity SaaSIdentity securitySREDevOpsPlatform engineering
Skills:MentorshipTechnical leadershipIncident reviewsDocumentationStandards setting
Languages:English
Tech Stack:AWSAzureApi gatewaysService meshesCachesConfiguration managementSecrets managementGitOpsTerraformOpenTofuAnsibleCI/CDDockerKubernetesGrafana CloudDatadogOpenTelemetryChaos engineeringDisaster recoverySLO

Company Brief

BeyondTrust
Provides privileged access management, vulnerability management, and secure remote access solutions to help organizations protect credentials, manage privileges, and secure endpoints across on-premises and cloud environments.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Funding: Private Equity Backed
Headquarters: Phoenix, United States
WebsiteLinkedIn