Site Reliability Engineering Manager

Onebrief
United States
Workplace: HybridFull timeUSD 205,000 - 255,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["People management","Incident response leadership","Communication","Risk assessment","Stakeholder management"]

Lead Onebrief’s SRE team within Infrastructure & Security to keep mission-critical deployments reliable, secure, and well supported across on-prem DoD and AWS environments. Manage engineers and work planning, coordinate delivery with platform, application, security, and customer success teams, and drive a reliability roadmap. Own operational readiness and incident response leadership (including blameless postmortems), improve observability and automation, and use SLIs/SLOs to reduce recurring failures and operational toil.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Onebrief
Onebrief
14 hours ago

Site Reliability Engineering Manager

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead Onebrief’s SRE team within Infrastructure & Security to keep mission-critical deployments reliable, secure, and well supported across on-prem DoD and AWS environments. Manage engineers and work planning, coordinate delivery with platform, application, security, and customer success teams, and drive a reliability roadmap. Own operational readiness and incident response leadership (including blameless postmortems), improve observability and automation, and use SLIs/SLOs to reduce recurring failures and operational toil.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Lead and develop the SRE team: hire, coach, manage performance, support career development, and build coverage across infrastructure and reliability.
  • •Own capacity and work planning: track incoming requests and support needs, break initiatives into tasks, sequence work, and adjust commitments as priorities/capacity change.
  • •Coordinate delivery across teams by managing dependencies with platform engineering, application engineering, security, and customer success; identify blockers early and communicate progress, risks, and decisions.
  • •Set and execute the reliability roadmap using customer needs, production data, incident patterns, and operational risks to prioritize improvements that reduce recurring failures.
  • •Establish operational ownership and lead incident response: define escalation paths and readiness criteria, coach incident commanders, run blameless postmortems/AARs, and assign follow-through for corrective actions.

Pay and Benefits

Salary: USD 205,000 - 255,000 annually
Equity and Bonus:Equity
Perks:Relocation

Key Requirements

  • •Active Secret clearance
  • •5+ years in Site Reliability Engineering, Platform Engineering, DevOps, or a related role with substantial infrastructure and operations experience
  • •Direct experience managing engineers, including coaching, performance management, career development, and hiring
  • •Experience planning team capacity, prioritizing competing requests, and coordinating engineering work across teams
  • •Experience with incident response and post-incident review practices, including helping engineers develop leadership and ownership
Experience:5+ yearsOn-prem environmentsAWS GovCloudDevOpsPlatform engineeringInfrastructure and operationsSRE
Skills:People managementIncident response leadershipCommunicationRisk assessmentStakeholder management
Certifications:AWS DevOps EngineerCKACKADSecurity+
Tech Stack:TerraformAnsiblePythonGoBashKubernetesCI/CDAWSAWS GovCloudGrafanaELKDatadogSLIsSLOsError budgetsGitOpsVMwareProxmoxNutanixHyper-V

Eligibility

Security Clearance:Secret

Company Brief

Onebrief
Provides an AI-enabled command operating system for military staff collaboration, operational planning, wargaming and decision support, used on classified and unclassified DoD networks to speed planning and execution.
Industry: Defense Technology
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Honolulu, United States
Founded: 2019
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor