Staff Site Reliability Engineer

Skydio
San Mateo
Workplace: RemoteFull timeUSD 240,000 - 300,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Problem-solving","Troubleshooting","Reliability focus"]

Own the cloud infrastructure that keeps the Skydio Cloud platform highly available for customers worldwide. Build, operate, and scale production Kubernetes/EKS clusters and AWS networking and security components using Terraform and CI/CD tooling. Triage and resolve production incidents across Kubernetes, AWS, Linux, networking, and databases, while improving observability, reliability, and scalability through automation and on-call rotations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Skydio
Skydio
1 day ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own the cloud infrastructure that keeps the Skydio Cloud platform highly available for customers worldwide. Build, operate, and scale production Kubernetes/EKS clusters and AWS networking and security components using Terraform and CI/CD tooling. Triage and resolve production incidents across Kubernetes, AWS, Linux, networking, and databases, while improving observability, reliability, and scalability through automation and on-call rotations.
Location: San Mateo
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, operate, and troubleshoot production Kubernetes/EKS clusters, including upgrades, node rollouts, and cluster maintenance.
  • •Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.
  • •Define and maintain infrastructure using Terraform and operate CI/CD and deployment infrastructure.
  • •Troubleshoot production incidents across Kubernetes, AWS, Linux, networking, and databases; improve monitoring, alerting, and observability.
  • •Participate in on-call rotations, respond to incidents, and identify and solve infrastructure scaling and reliability problems, automating operational work with Python/Go.

Pay and Benefits

Salary: USD 240,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceEquity401kPaid LeaveHoliday Pay

Key Requirements

  • •8+ years in Site Reliability Engineering, Platform Engineering, DevOps, Production Engineering, or an equivalent infrastructure role.
  • •Strong, hands-on experience operating Kubernetes (not just deploying to existing clusters), including managing Kubernetes/EKS upgrades and production clusters.
  • •Strong AWS fundamentals including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.
  • •Production experience with Terraform (or similar infrastructure-as-code).
  • •Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.
Experience:8+ yearsDronesAutonomous flightCloud infrastructure
Skills:Problem-solvingTroubleshootingReliability focus
Tech Stack:KubernetesAmazon EKSAWSTerraformCI/CDObservabilityNetworkingLinuxPythonGoVPCSubnetsLoad balancersIAMDatabasesArgo CDSpinnakerGitHub ActionsGitLab CI/CDJenkins

Company Brief

Skydio
Designs and manufactures autonomous consumer and enterprise drones using AI-powered vision and navigation systems for inspection, mapping, public safety, and industrial applications. Focuses on rugged, fully autonomous flight and end-to-end drone solutions.
Industry: Drones & Aerial Logistics
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Mateo, United States
Founded: 2014
WebsiteLinkedIn