Sr. Staff Lead Site Reliability Engineer (R5803)

Shield AI
San Mateo
Workplace: OnsiteFull timeUSD 220,000 - 330,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Root-cause analysis","Technical leadership","Mentoring","Capacity planning","Thought leadership"]

Own and mature reliability engineering practices for Hivemind’s cloud infrastructure and platform services. Define SLIs/SLOs, build observability (monitoring, alerting, logging, tracing), and lead complex incident response with root-cause analysis. Partner with product and Cloud Engineering to bake reliability into system design, improve resilience via automation/testing/capacity planning, and develop operational tooling. Mentor and lead adoption of reliability-first operational practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Shield AI
Shield AI
2 days ago

Sr. Staff Lead Site Reliability Engineer (R5803)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Own and mature reliability engineering practices for Hivemind’s cloud infrastructure and platform services. Define SLIs/SLOs, build observability (monitoring, alerting, logging, tracing), and lead complex incident response with root-cause analysis. Partner with product and Cloud Engineering to bake reliability into system design, improve resilience via automation/testing/capacity planning, and develop operational tooling. Mentor and lead adoption of reliability-first operational practices.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define and implement SLIs, SLOs, and other service reliability measures.
  • •Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
  • •Lead technical response to complex incidents and drive root-cause analysis through resolution.
  • •Improve system resilience via automation, testing, capacity planning, and failure recovery; reduce manual operational work with tooling.
  • •Establish incident response practices and partner with product/platform teams to incorporate reliability requirements into system design; mentor teammates and manage a short- and long-term SRE roadmap.

Pay and Benefits

Salary: USD 220,000 - 330,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
  • •Experience operating production services with defined availability and reliability requirements.
  • •Experience implementing SLIs/SLOs, monitoring, alerting, and incident response practices.
  • •Experience designing and operating infrastructure in AWS or another major cloud environment.
  • •Experience implementing infrastructure-as-code, supporting containerized apps/distributed systems, and building operational tooling/automation (e.g., Python, Go).
Experience:Cloud-nativeDistributed systemsProduction services
Skills:Root-cause analysisTechnical leadershipMentoringCapacity planningThought leadership
Tech Stack:AWSPythonGoKubernetes

Company Brief

Shield AI
Develops AI-powered autonomy and software for military aircraft and drones to enable autonomous ISR and combat missions, integrating perception, navigation, and mission planning for defense customers.
Industry: Defense Technology
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series E+
Headquarters: San Diego, United States
Founded: 2015
WebsiteLinkedIn