Senior AI Infrastructure Engineer

Anduril
Seattle, Washington
Workplace: OnsiteFull timeUSD 191,000 - 253,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Design reviews","Code review","Mentorship","Engineering best practices","Performance optimization"]

Build and operate end-to-end infrastructure for training, evaluating, hosting, and serving advanced AI models powering autonomous systems. Own critical ML platform components and MLOps tooling across cloud and air-gapped edge environments. Develop scalable training and experimentation pipelines, multi-modal ETL, and high-throughput low-latency model serving, plus CI/CD, evaluation/validation, and safe rollout. Partner with AI Research and platform engineers and mentor peers through design and code reviews.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anduril
Anduril
19 hours ago

Senior AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build and operate end-to-end infrastructure for training, evaluating, hosting, and serving advanced AI models powering autonomous systems. Own critical ML platform components and MLOps tooling across cloud and air-gapped edge environments. Develop scalable training and experimentation pipelines, multi-modal ETL, and high-throughput low-latency model serving, plus CI/CD, evaluation/validation, and safe rollout. Partner with AI Research and platform engineers and mentor peers through design and code reviews.
Location: Seattle, Washington
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate model development.
  • •Identify and resolve ML lifecycle bottlenecks by creating tooling for experiment tracking, automated profiling, and hyperparameter tuning.
  • •Implement and scale robust multi-modal data pipelines to process video feeds, radar, flight telemetry, and simulation logs.
  • •Deploy high-throughput, low-latency model serving frameworks for both cloud and resource-constrained, air-gapped edge networks.
  • •Develop CI/CD and evaluation pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies.

Pay and Benefits

Salary: USD 191,000 - 253,000 annually

Key Requirements

  • •5+ years of software engineering experience building and operating production-scale machine learning infrastructure or distributed systems.
  • •Proficiency in Python, Go, or C++ with strong software engineering fundamentals, systems design, and concurrent programming.
  • •Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, Megatron-LM).
  • •Experience building and maintaining distributed data pipelines for large-scale unstructured or multi-modal datasets.
  • •Eligible to obtain and maintain an active U.S. Top Secret security clearance.
Experience:5+ yearsMachine learningDistributed systems
Skills:Design reviewsCode reviewMentorshipEngineering best practicesPerformance optimization
Tech Stack:PythonGoC++DockerKubernetesPyTorch DistributedRaySlurmMegatron-LMETLCI/CDLLMsComputer visionRL agentsRLHFDPOSLAMData pipelinesDistributed trainingExperiment tracking

Eligibility

Security Clearance:Top Secret

Company Brief

Anduril
Designs and builds advanced defense systems combining autonomous aircraft, sensors, and AI-driven software for military and national security applications, focused on modernizing battlefield capabilities and distributed sensing.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Costa Mesa, United States
Founded: 2017
WebsiteLinkedIn