Senior Site Reliability Engineer, Application Reliability (Arlington, VA) - Secret Clearance Required - Relocation Provided

Onebrief
United States
Workplace: HybridFull timeUSD 180,000 - 220,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Mentorship","Blameless postmortems","Incident response","Root cause analysis"]

Join the Infrastructure & Security team as a Site Reliability Engineer focused on application reliability. Work primarily in TypeScript to fix performance and reliability issues at the source, build observability (Prometheus/Loki/Alloy/Grafana), define SLIs/SLOs, and lead incident response with blameless postmortems. You’ll collaborate closely with product, security, and customer success while improving mission-critical deployments across on-prem DoD and AWS environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Onebrief
Onebrief
1 month ago

Senior Site Reliability Engineer, Application Reliability (Arlington, VA) - Secret Clearance Required - Relocation Provided

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Join the Infrastructure & Security team as a Site Reliability Engineer focused on application reliability. Work primarily in TypeScript to fix performance and reliability issues at the source, build observability (Prometheus/Loki/Alloy/Grafana), define SLIs/SLOs, and lead incident response with blameless postmortems. You’ll collaborate closely with product, security, and customer success while improving mission-critical deployments across on-prem DoD and AWS environments.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Improve production reliability and performance by working in the application codebase (primarily TypeScript) and partnering with product engineers on design and reviews.
  • •Design and run monitoring, logging, and alerting using Prometheus, Loki, Alloy, and Grafana so teams catch issues based on real application behavior.
  • •Define and measure SLIs and SLOs and wire alerting to support reliability targets.
  • •Act as incident responder/incident commander when needed and run blameless postmortems to drive code or process fixes.
  • •Automate away repetitive operational work by writing software to reduce toil and helping teams get production-ready, including in air-gapped environments.

Pay and Benefits

Salary: USD 180,000 - 220,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Active Secret clearance.
  • •5+ years in software engineering, SRE, or a related role with real time writing and shipping application code.
  • •Strong TypeScript experience (or comparable modern language experience with willingness to work primarily in TypeScript).
  • •Solid grasp of the full SDLC (design, code review, testing, release) and how reliability fits each stage.
  • •Experience with incident response, root cause analysis, and turning findings into lasting fixes.
Skills:CollaborationMentorshipBlameless postmortemsIncident responseRoot cause analysis
Certifications:AWS DevOps EngineerCKACKAD
Tech Stack:TypeScriptNodeKubectlPrometheusLokiAlloyGrafanaSLIsSLOsGitHub ActionsGitLab CI/CDJenkinsPythonGoBashKubernetesAWSAWS GovCloudELKDatadog

Eligibility

Security Clearance:Secret

Company Brief

Onebrief
Provides an AI-enabled command operating system for military staff collaboration, operational planning, wargaming and decision support, used on classified and unclassified DoD networks to speed planning and execution.
Industry: Defense Technology
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Honolulu, United States
Founded: 2019
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor