Site Reliability Engineer

GovCIO
United States
Workplace: HybridFull timeUSD 230,000 - 250,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: []

Ensure reliability, scalability, performance, and availability of mission-critical systems by designing and maintaining highly available production environments. Define and manage SLI/SLO targets and error budgets, automate operational work, and build monitoring, alerting, and observability capabilities. Lead incident response and root-cause analysis, improve performance and resilience, and implement disaster recovery and continuity strategies while partnering with development teams to strengthen application reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
GovCIO
GovCIO
1 month ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 37 minutes agoStatus: Live

Job Summary

Ensure reliability, scalability, performance, and availability of mission-critical systems by designing and maintaining highly available production environments. Define and manage SLI/SLO targets and error budgets, automate operational work, and build monitoring, alerting, and observability capabilities. Lead incident response and root-cause analysis, improve performance and resilience, and implement disaster recovery and continuity strategies while partnering with development teams to strengthen application reliability.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and maintain highly available production systems.
  • •Define and manage SLIs, SLOs, and error budgets.
  • •Automate operational tasks and eliminate manual processes.
  • •Develop monitoring, alerting, and observability solutions to improve performance, capacity, and resilience.
  • •Lead incident response and root cause analysis, including implementing disaster recovery and continuity strategies.

Pay and Benefits

Salary: USD 230,000 - 250,000 annually

Key Requirements

  • •12+ years of infrastructure/cloud engineering experience (or commensurate experience).
  • •5–10+ years of engineering experience with strong Linux and Windows systems background.
  • •Expertise in Kubernetes and container platforms.
  • •Proficiency in scripting languages such as Python and Go; hands-on knowledge of Terraform and automation tools.
  • •Experience designing and managing CI/CD pipelines and monitoring/incident management practices.
Experience:5+ yearsInfrastructureCloud engineeringDevOpsSRE
Education:Bachelor's
Certifications:Kubernetes certificationsAWS/Azure certificationsDevOps certificationsITIL
Languages:English
Tech Stack:LinuxWindowsKubernetesContainer platformsPythonGoTerraformAutomationMonitoring platformsCI/CD pipelinesAWSAzure

Eligibility

Security Clearance:Secret

Company Brief

GovCIO
Provides IT modernization, cloud, cybersecurity, data analytics, and managed services to U.S. federal civilian and defense agencies, delivering technology solutions and mission support to improve government operations and security.
Industry: Professional Services
Growth: Established Company
Headquarters: Herndon, United States
WebsiteLinkedIn