Site Reliability Engineering Lead

Truist Financial
Atlanta, Raleigh, Charlotte
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Leadership","Mentoring","Incident management","Communication"]

Lead enterprise reliability for hybrid cloud and on-premises platforms, driving automation, observability, and incident/problem management improvements. Standardize SRE practices across teams, establish incident playbooks and SLO/SLI adoption, and design scalable, secure, highly available solutions. Mentor SRE engineers, evaluate emerging technologies, and deliver measurable reductions in downtime and MTTR using tools and cloud-native operational patterns.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Truist Financial
Truist Financial
58 minutes ago

Site Reliability Engineering Lead

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 58 minutes agoStatus: Live

Job Summary

Lead enterprise reliability for hybrid cloud and on-premises platforms, driving automation, observability, and incident/problem management improvements. Standardize SRE practices across teams, establish incident playbooks and SLO/SLI adoption, and design scalable, secure, highly available solutions. Mentor SRE engineers, evaluate emerging technologies, and deliver measurable reductions in downtime and MTTR using tools and cloud-native operational patterns.
Location: Atlanta, Raleigh, Charlotte
Workplace: Onsite
Employment Type: Full time · Permanent
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead major and high-severity incident response, driving multi-team diagnosis and technical resolution and driving problem management to closure.
  • •Architect and deliver automation to eliminate toil, reduce MTTR, and improve service resilience.
  • •Enhance observability across logs, metrics, traces, and events, and standardize enterprise observability practices, dashboards, and KPIs.
  • •Establish and maintain incident playbooks, escalation paths, and communication frameworks; guide SLO/SLI adoption across product teams.
  • •Mentor and coach SRE engineers, lead workshops and enablement, and evaluate emerging technologies by building prototypes and solution concepts.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife InsuranceDisability401kPaid Leave

Key Requirements

  • •Bachelor’s degree in Computer Science, Software Engineering, or a related field.
  • •Minimum 7 years of professional experience in software development.
  • •Deep knowledge of software architecture and design principles across multiple programming languages.
  • •Deep understanding of the software development lifecycle, testing, deployment, and security practices.
  • •Experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations (7+ years), including distributed systems and Kubernetes.
Education:Bachelor's
Skills:LeadershipMentoringIncident managementCommunication
Languages:English
Tech Stack:KubernetesPythonGoPowerShellAnsibleSplunkDynatraceAIAIOpsCI/CDMicroservicesLinux/UnixObservability

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Truist Financial
Provides consumer and commercial banking, wealth management, insurance, lending, and payments services through a large U.S. financial services platform formed by the merger of BB&T and SunTrust.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Charlotte, United States
Founded: 2019
WebsiteLinkedIn