Digital - Principal SRE

Huntington Bancshares
Columbus
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Attention to detail","Collaboration","Technical documentation"]

Design, deploy, and maintain AI-driven solutions while applying SRE best practices to ensure reliability, scalability, and performance of digital platforms. Monitor AI application and infrastructure health, define SLOs/error budgets, and lead incident response for AI services. Integrate machine learning models into production, build AI platform integration across providers, and drive continuous improvement through automation, observability, and performance optimization.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Huntington Bancshares
Huntington Bancshares
2 months ago

Digital - Principal SRE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design, deploy, and maintain AI-driven solutions while applying SRE best practices to ensure reliability, scalability, and performance of digital platforms. Monitor AI application and infrastructure health, define SLOs/error budgets, and lead incident response for AI services. Integrate machine learning models into production, build AI platform integration across providers, and drive continuous improvement through automation, observability, and performance optimization.
Location: Columbus
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design, develop, and implement AI-driven systems and automation tools to improve reliability and efficiency of digital platforms.
  • •Monitor AI-enabled application and infrastructure health, availability, and performance using SRE best practices.
  • •Integrate machine learning models into production environments to enable seamless deployment and operation.
  • •Establish and enforce SLOs, error budgets, and incident response procedures for AI-driven services.
  • •Troubleshoot complex AI-related incidents using observability/monitoring, and drive continuous improvement through automation and optimization.

Key Requirements

  • •Bachelor’s or master’s degree in computer science, engineering, data science, or a related field.
  • •Minimum 5 years of proven experience in AI/ML engineering, SRE, DevOps, or related roles.
  • •Strong programming skills in Python, Java, or similar languages with experience developing and deploying machine learning models.
  • •Hands-on experience with cloud platforms (AWS, GCP, Azure) and containerization (Docker, Kubernetes).
  • •Experience with observability (Prometheus, Grafana, ELK stack), ServiceNow incident management, infrastructure-as-code (Terraform, Ansible), and CI/CD pipelines.
Experience:5+ yearsAI/MLSREDevOpsCloud-native AI
Education:Bachelor's
Skills:Problem-solvingAttention to detailCollaborationTechnical documentation
Tech Stack:PythonJavaAWSGCPAzureDockerKubernetesPrometheusGrafanaELKServiceNowTerraformAnsibleCI/CDOpenAIGoogle

Company Brief

Huntington Bancshares
Regional bank holding company providing commercial and consumer banking, payments, wealth management, and lending services across the Midwestern and select other U.S. markets. It serves individuals, small businesses, and corporate clients through branches and digital channels.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Columbus, United States
Founded: 1866
WebsiteLinkedIn