Senior Staff Lead Site Reliability Engineer (R5803)

Shield AI
San Diego
Workplace: OnsiteFull timeUSD 183,000 - 275,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Technical leadership","Mentorship","Root-cause analysis","Incident response","Diagnosing complex failures"]

Establish and mature reliability practices across cloud infrastructure and platform services, defining SLO/SLI targets and improving observability. Build and enhance monitoring, alerting, logging, and tracing, lead technical response to complex incidents, and drive root-cause analysis. Identify recurring failure modes, improve resilience via automation/testing/capacity planning, and develop operational tooling. Mentor engineers and partner with product and platform teams to embed reliability into system design.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Shield AI
Shield AI
2 days ago

Senior Staff Lead Site Reliability Engineer (R5803)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Establish and mature reliability practices across cloud infrastructure and platform services, defining SLO/SLI targets and improving observability. Build and enhance monitoring, alerting, logging, and tracing, lead technical response to complex incidents, and drive root-cause analysis. Identify recurring failure modes, improve resilience via automation/testing/capacity planning, and develop operational tooling. Mentor engineers and partner with product and platform teams to embed reliability into system design.
Location: San Diego
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and implement SLIs, SLOs, and other service reliability measures.
  • •Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
  • •Lead technical response to complex incidents and drive root-cause analysis through resolution.
  • •Improve system resilience through automation, testing, capacity planning, and failure recovery.
  • •Establish incident response practices and mentor teammates in reliability and operational practices.

Pay and Benefits

Salary: USD 183,000 - 275,000 annually
Equity and Bonus:Equity

Key Requirements

  • •7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
  • •Experience operating production services with defined availability and reliability requirements.
  • •Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
  • •Experience designing and operating infrastructure in AWS or another major cloud environment.
  • •Experience with infrastructure-as-code and automated infrastructure provisioning.
Experience:7+ yearsSRECloud infrastructureInfrastructure engineeringPlatform services
Skills:Technical leadershipMentorshipRoot-cause analysisIncident responseDiagnosing complex failures
Tech Stack:AWSPythonGoKubernetesSLIsSLOsMonitoringAlertingLoggingTracingInfrastructure as codeContainerized applicationsDistributed systems

Company Brief

Shield AI
Develops AI-powered autonomy and software for military aircraft and drones to enable autonomous ISR and combat missions, integrating perception, navigation, and mission planning for defense customers.
Industry: Defense Technology
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series E+
Headquarters: San Diego, United States
Founded: 2015
WebsiteLinkedIn