Technical Site Reliability Engineer

Anduril
Abu Dhabi, London
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Troubleshooting","Issue triage","Documentation","Cross-team collaboration"]

Design, build, and operate the infrastructure that keeps large-scale military simulation facilities running reliably. Maintain the simulation software stack, provision and patch compute/networking/storage, and build automated post-release regression and smoke tests to catch silent regressions early. Diagnose failure modes, monitor environment health, partner with development teams to mitigate reliability risks, and document runbooks and validation results for the Simulation Center.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
11 hours ago

Technical Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design, build, and operate the infrastructure that keeps large-scale military simulation facilities running reliably. Maintain the simulation software stack, provision and patch compute/networking/storage, and build automated post-release regression and smoke tests to catch silent regressions early. Diagnose failure modes, monitor environment health, partner with development teams to mitigate reliability risks, and document runbooks and validation results for the Simulation Center.
Location: Abu Dhabi, London
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Maintain the simulation software stack (installation, configuration, updates, version management, and day-to-day functionality).
  • •Own and operate the underlying infrastructure (compute, networking, storage, and environment configuration) to keep it provisioned, patched, and performant.
  • •Build and maintain a post-release regression and smoke-test suite that runs after every release or configuration change.
  • •Forecast, diagnose, and eliminate failure modes via root-cause analysis and guardrails/monitoring to prevent recurrence.
  • •Partner with development teams to review changes for reliability risk and implement mitigation strategies before releases land.

Pay and Benefits

Perks:Health InsuranceEquity Grants

Key Requirements

  • •Proficiency in Python for automation, tooling, and test development.
  • •Working knowledge of C++ to read, debug, build, and trace issues in the simulation codebase.
  • •Solid networking fundamentals (TCP/IP, UDP, multicast, DNS, routing, firewalls) and ability to diagnose latency, packet loss, and connectivity issues.
  • •Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
  • •Demonstrated experience maintaining production (or production-adjacent) systems, including troubleshooting under time pressure, plus strong written and verbal communication.
Experience:DefenseSimulationWargaming
Skills:CommunicationTroubleshootingIssue triageDocumentationCross-team collaboration
Tech Stack:PythonC++TCP/IPUDPMulticastDNSRoutingFirewallsTerraformAnsibleDockerKubernetesPrometheusGrafanaELKLinuxWindowsDISHLATENA

Eligibility

Security Clearance:Active security clearance

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn