Site Reliability Engineer (SRE) – II

Huntington Bancshares
United States
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Problem-solving","Troubleshooting","Mentorship","Communication","Documentation"]

Maintain the availability, scalability, and performance of critical services as an SRE Level II. Build and automate reliability improvements using infrastructure as code, lead incident troubleshooting with detailed RCA, and participate in on-call escalations. Partner with software teams on resilient system design, capacity planning, and performance optimization. Improve monitoring and observability, support security and compliance, and mentor junior SREs while driving continuous operational improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Huntington Bancshares
Huntington Bancshares
2 months ago

Site Reliability Engineer (SRE) – II

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Maintain the availability, scalability, and performance of critical services as an SRE Level II. Build and automate reliability improvements using infrastructure as code, lead incident troubleshooting with detailed RCA, and participate in on-call escalations. Partner with software teams on resilient system design, capacity planning, and performance optimization. Improve monitoring and observability, support security and compliance, and mentor junior SREs while driving continuous operational improvements.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Maintain service uptime by managing complex incidents and ensuring reliability during high-impact production issues with RCA and preventative measures.
  • •Develop and maintain automation scripts and infrastructure as code using tools such as Terraform, Ansible, or CloudFormation.
  • •Analyze performance metrics and recommend optimizations for scalability and reliability, including support for capacity planning.
  • •Collaborate with software engineering teams on scalable, resilient system design and architecture decisions for fault tolerance, redundancy, and recovery.
  • •Build and optimize monitoring, alerting, and observability solutions using tools like Prometheus, Grafana, Datadog, and Dynatrace to detect and resolve issues proactively.

Key Requirements

  • •Minimum 5 years of experience in site reliability engineering, DevOps, systems administration, or related roles.
  • •Strong Linux/Unix administration and scripting proficiency (Python, Bash, or Go).
  • •Deep understanding of cloud platforms (AWS, GCP, Azure) and services such as EC2, S3, Lambda, and Kubernetes.
  • •Experience with containerization/orchestration technologies like Docker and Kubernetes.
  • •Proficiency with monitoring/observability tools such as Dynatrace, Prometheus, Grafana, Datadog, or ELK Stack, plus CI/CD and infrastructure automation tools.
Experience:5+ years
Skills:Problem-solvingTroubleshootingMentorshipCommunicationDocumentation
Tech Stack:LinuxUnixPythonBashGoTerraformAnsibleCloudFormationAWSGCPAzureEC2S3LambdaKubernetesDockerPrometheusGrafanaDatadogDynatrace

Company Brief

Huntington Bancshares
Regional bank holding company providing commercial and consumer banking, payments, wealth management, and lending services across the Midwestern and select other U.S. markets. It serves individuals, small businesses, and corporate clients through branches and digital channels.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Columbus, United States
Founded: 1866
WebsiteLinkedIn