Staff Site Reliability Engineer

Anduril
United States
Workplace: OnsiteFull timeUSD 191,000 - 253,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Written communication","Verbal communication","Independent problem-solving","Influencing senior engineers","Cross-functional collaboration"]

Own reliability architecture for systems running Anduril’s business and manufacturing operations, shaping how observability, deployments, incident management, and capacity planning work at platform scale. Design and operate metrics, tracing, logging, alerting, and dashboards; build release-safety mechanisms like progressive rollouts, canary analysis, and rollback gates. Define SLO and production-readiness frameworks, lead complex incident response, and help teams embed reliability across the engineering lifecycle.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
8 hours ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own reliability architecture for systems running Anduril’s business and manufacturing operations, shaping how observability, deployments, incident management, and capacity planning work at platform scale. Design and operate metrics, tracing, logging, alerting, and dashboards; build release-safety mechanisms like progressive rollouts, canary analysis, and rollback gates. Define SLO and production-readiness frameworks, lead complex incident response, and help teams embed reliability across the engineering lifecycle.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Set the reliability architecture for CorpTech Platform’s production environment, including observability infrastructure, deployment systems, incident management, capacity planning, and failure-domain isolation.
  • •Design and operate the observability platform with metrics, distributed tracing, structured logging, alerting, and dashboarding at scale.
  • •Own deployment and release-safety mechanisms such as progressive rollouts, canary analysis, automated rollback, and deployment gates.
  • •Define and govern SLO frameworks and production-readiness standards, embedding reliability into design and pre-launch processes.
  • •Lead incident response for complex multi-system failures and drive post-incident improvements; develop reliability patterns for AI-enabled systems and capacity/cost/performance engineering for critical paths.

Pay and Benefits

Salary: USD 191,000 - 253,000 annually
Equity and Bonus:Equity

Key Requirements

  • •10+ years of experience in site reliability engineering, production engineering, infrastructure engineering, or a closely related discipline, including architecture/platform-wide scope.
  • •Experience designing and owning reliability infrastructure that multiple engineering teams depend on, such as observability platforms, deployment systems, or incident management tooling at meaningful scale.
  • •Deep technical fluency across distributed systems, Kubernetes, cloud platforms (AWS, GCP, or Azure), networking, storage, and layer-specific failure modes.
  • •Proficiency in systems programming languages (Go, Python, Rust, or equivalent) for production infrastructure and automation tooling.
  • •Demonstrated ability to define SRE standards or production-readiness frameworks and influence adoption across engineering teams without formal authority.
Experience:10+ yearsSite reliability engineering
Education:Bachelor's
Skills:Written communicationVerbal communicationIndependent problem-solvingInfluencing senior engineersCross-functional collaboration
Languages:English
Tech Stack:KubernetesAWSGCPAzureGoPythonRustDatadogGrafanaPrometheusOpenTelemetry

Eligibility

Security Clearance:U.S. Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn