Site Reliability Engineer

Anduril
Massachusetts
Workplace: OnsiteFull timeUSD 166,000 - 220,000Function: DevOps, Cloud & InfrastructureSkills: ["Ownership","Calm under pressure","Proactive communication","Methodical troubleshooting","Documentation"]

Own the health and uptime of deployed imaging systems, triaging and diagnosing issues across networks, services, calibration, upgrades, and sensor hardware. Serve as the frontline for escalations from field personnel and support channels, turning recurring problems into runbooks, diagnostics, and self-service tooling. When defects are code-level, reproduce, document, and hand off to Mission Software Engineers while feeding reliability learnings back into product observability and upgrade safety.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
2 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 minutes agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Own the health and uptime of deployed imaging systems, triaging and diagnosing issues across networks, services, calibration, upgrades, and sensor hardware. Serve as the frontline for escalations from field personnel and support channels, turning recurring problems into runbooks, diagnostics, and self-service tooling. When defects are code-level, reproduce, document, and hand off to Mission Software Engineers while feeding reliability learnings back into product observability and upgrade safety.
Location: Massachusetts
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own fielded system reliability: ensure health and uptime and drive resolution when issues arise
  • •Run point on escalations from support channels and customer-support pipeline (Tier 0 through escalation paths)
  • •Build and maintain runbooks, diagnostics, and self-service tooling to reduce repeat incidents
  • •Hold the boundary with engineering: reproduce and document software defects and hand off to Mission Software Engineers
  • •Turn field learnings into reliability improvements for observability, safer upgrades, and more graceful failure
Travel: Medium travel

Pay and Benefits

Salary: USD 166,000 - 220,000
Equity and Bonus:Equity

Key Requirements

  • •3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems with ownership after systems ship
  • •Strong Linux fundamentals, including troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained/field environments)
  • •Ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without full visibility into every component
  • •Own a structured on-call rotation, including scheduled after-hours and weekend coverage
  • •Strong written and verbal communication to run remote troubleshooting with non-technical operators and document outcomes
Experience:SREDevOpsField/systems engineeringProduction supportDeployed hardware/software
Skills:OwnershipCalm under pressureProactive communicationMethodical troubleshootingDocumentation
Languages:English
Tech Stack:LinuxIPRoutingVPNPythonBashNixNixOSSystemdObservabilityPagerDutyLattice OS

Eligibility

Security Clearance:Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn