Staff DevOps Engineer

Anduril
United States
Workplace: OnsiteFull timeUSD 191,000 - 253,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Written and verbal communication","Influencing senior engineers","Independent problem-solving","Working through ambiguity","Leadership in incident response"]

Set the reliability architecture for production environments across CorpTech Platform, owning observability, deployment/release-safety, incident management, and capacity planning. Design and operate large-scale monitoring with metrics, distributed tracing, structured logging, alerting, and dashboards. Define SLO frameworks and production-readiness standards, identify systemic reliability risks, lead complex incident response, and drive durable infrastructure improvements aligned with AI-enabled system reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
1 day ago

Staff DevOps Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Set the reliability architecture for production environments across CorpTech Platform, owning observability, deployment/release-safety, incident management, and capacity planning. Design and operate large-scale monitoring with metrics, distributed tracing, structured logging, alerting, and dashboards. Define SLO frameworks and production-readiness standards, identify systemic reliability risks, lead complex incident response, and drive durable infrastructure improvements aligned with AI-enabled system reliability.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Set reliability architecture for CorpTech Platform production, including observability infrastructure, deployment systems, incident management, capacity planning, and failure-domain isolation.
  • •Design and operate observability platforms (metrics, distributed tracing, structured logging, alerting, and dashboarding) for production visibility at scale.
  • •Own deployment infrastructure and release-safety mechanisms, including progressive rollout systems, canary analysis, automated rollback, and deployment gates.
  • •Define and govern SLO frameworks and production-readiness standards that teams adopt during design and pre-launch, embedding reliability into the development lifecycle.
  • •Lead incident response for complex, multi-system failures and drive post-incident processes that produce durable systemic improvements.

Pay and Benefits

Salary: USD 191,000 - 253,000 annually

Key Requirements

  • •10+ years of experience in site reliability engineering, production engineering, infrastructure engineering, or related fields with architecture/platform-wide scope.
  • •Experience designing and owning reliability infrastructure used by multiple engineering teams (e.g., observability platforms, deployment systems, incident management tooling).
  • •Deep fluency across distributed systems, Kubernetes/container orchestration, cloud platforms (AWS, GCP, or Azure), networking, storage, and layer-specific failure modes.
  • •Proficiency in systems programming languages (Go, Python, Rust, or equivalent) for infrastructure tooling and automation.
  • •Experience defining SRE standards and production-readiness frameworks and influencing adoption without formal authority.
Experience:10+ yearsSite reliability engineeringProduction engineeringInfrastructure engineeringDistributed systemsCloud
Education:Bachelor's in Computer Science, Information Systems, Engineering, or related technical field
Skills:Written and verbal communicationInfluencing senior engineersIndependent problem-solvingWorking through ambiguityLeadership in incident response
Tech Stack:KubernetesAWSGCPAzureGoPythonRustObservabilityDistributed tracingStructured loggingAlertingDashboardingSLOCanary analysisProgressive rolloutAutomated rollbackOpenTelemetryPrometheusGrafanaDatadog

Eligibility

Security Clearance:Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn