Senior Infrastructure Reliability Engineer

Anduril
United States
Workplace: OnsiteFull timeUSD 166,000 - 220,000 annuallyFunction: Solutions Engineering & Sales EngineeringSkills: ["Continuous improvement","Automation mindset","Problem-solving","Cross-functional collaboration","End-to-end ownership"]

Own the lifecycle of critical self-hosted developer infrastructure—from patching and upgrades to backups, scaling, and incident response. You’ll run and troubleshoot platforms used daily by engineers, using SRE practices and automation to improve reliability. Build Infrastructure-as-Code with Terraform, define SLOs, and lead root-cause analysis. Work across platform, security, on-prem/cloud infrastructure, and software teams while operating Kubernetes and containerized systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
2 days ago

Senior Infrastructure Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own the lifecycle of critical self-hosted developer infrastructure—from patching and upgrades to backups, scaling, and incident response. You’ll run and troubleshoot platforms used daily by engineers, using SRE practices and automation to improve reliability. Build Infrastructure-as-Code with Terraform, define SLOs, and lead root-cause analysis. Work across platform, security, on-prem/cloud infrastructure, and software teams while operating Kubernetes and containerized systems.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team.
  • •Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray).
  • •Design and implement automated systems for patching, backups (with validation), and upgrades.
  • •Scale infrastructure for a fast-growing engineering organization and manage environments with Infrastructure-as-Code (Terraform).
  • •Define and maintain SLOs, build monitoring/alerting/observability, and lead incident response and root cause analysis.

Pay and Benefits

Salary: USD 166,000 - 220,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Operate infrastructure outside managed cloud services, including bare-metal Kubernetes and on-prem virtualization (VMware ESXi/vSphere).
  • •Run production systems using Docker and Kubernetes.
  • •Have strong Linux foundations (RHEL, Ubuntu).
  • •Be proficient with at least one cloud platform (AWS, GCP, or Azure).
  • •Manage infrastructure with Infrastructure-as-Code (e.g., Terraform/OpenTofu) and configuration management tools (e.g., Ansible, Puppet, Chef), plus automation and scripting (Python, Go, Bash).
Experience:SREDevOpsOn-prem infrastructureKubernetesInfrastructure automationCI/CD
Skills:Continuous improvementAutomation mindsetProblem-solvingCross-functional collaborationEnd-to-end ownership
Tech Stack:RunAIGitHub Enterprise ServerCircleCIJFrog ArtifactoryXrayTerraformOpenTofuAnsiblePuppetChefPythonGoBashLinuxRHELUbuntuDockerKubernetesVMware ESXiVSphere

Eligibility

Security Clearance:Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn