Staff Site Reliability Engineer

Anduril
United States
Workplace: OnsiteFull timeUSD 191,000 - 253,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Reliability architecture","Incident response","Stakeholder influence","Communication","Independent decision-making"]

Own the reliability architecture for systems powering business and manufacturing operations. Build and operate observability (metrics, distributed tracing, structured logging), deployment and release-safety mechanisms (progressive rollout, canary, rollback, deployment gates), and SLO frameworks that translate trade-offs into measurable outcomes. Lead incident response for complex multi-system failures, drive systemic reliability investments, and design reliability patterns for AI-enabled workloads across the platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
1 day ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 44 minutes agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Own the reliability architecture for systems powering business and manufacturing operations. Build and operate observability (metrics, distributed tracing, structured logging), deployment and release-safety mechanisms (progressive rollout, canary, rollback, deployment gates), and SLO frameworks that translate trade-offs into measurable outcomes. Lead incident response for complex multi-system failures, drive systemic reliability investments, and design reliability patterns for AI-enabled workloads across the platform.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Set reliability architecture for the production environment, including observability, deployment systems, incident management, capacity planning, and failure-domain isolation.
  • •Design and operate the observability platform (metrics, distributed tracing, structured logging, alerting, dashboarding) for large-scale production visibility.
  • •Own deployment and release-safety mechanisms, including progressive rollout, canary analysis, automated rollback, and deployment gates.
  • •Define and govern SLO frameworks to make reliability measurable and actionable across SRE, product engineering, and leadership.
  • •Lead incident response for complex multi-system failures and drive post-incident processes that result in durable systemic improvements.

Pay and Benefits

Salary: USD 191,000 - 253,000 annually
Equity and Bonus:Equity

Key Requirements

  • •10+ years of site reliability, production, infrastructure, or closely related experience with architecture or platform-wide scope.
  • •Experience designing and owning reliability infrastructure that multiple engineering teams depend on (observability, deployment systems, incident management) at meaningful scale.
  • •Deep technical fluency across distributed systems, Kubernetes, cloud platforms (AWS/GCP/Azure), networking, storage, and layer-specific failure modes.
  • •Proficiency in systems programming languages (Go, Python, Rust or equivalent) for production infrastructure, tooling, and automation.
  • •Demonstrated experience defining SRE standards and production-readiness frameworks, influencing adoption without formal authority; experience leading complex incident response and turning findings into infrastructure investments.
Experience:10+ yearsDistributed systemsCloud infrastructureObservabilitySite reliability engineeringAI-enabled systems
Education:Bachelor's in Computer Science, Information Systems, Engineering, or related technical field
Skills:Reliability architectureIncident responseStakeholder influenceCommunicationIndependent decision-making
Tech Stack:KubernetesAWSGCPAzureGoPythonRustDatadogGrafanaPrometheusOpenTelemetry

Eligibility

Security Clearance:Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn