Associate Director, Software Engineering (AI Platforms SRE) (Guangzhou, GD, CN, 510620)

HSBC
Guangzhou
Full timeFunction: Software EngineeringSkills: ["Problem-solving","Analytical","Communication","Mentoring"]

Lead SRE for AI platforms within the CTO Platforms business, driving incident troubleshooting and root-cause prevention while designing resilient, secure infrastructure across cloud and container ecosystems. Own SRE practices such as SLIs/SLOs, error budgets, and automation to reduce toil. Build observability and deployment automation (CI/CD, infrastructure-as-code), mentor engineers, and partner across engineering, QA, product, and operations to embed reliability end-to-end.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
HSBC
HSBC
20 hours ago

Associate Director, Software Engineering (AI Platforms SRE) (Guangzhou, GD, CN, 510620)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Lead SRE for AI platforms within the CTO Platforms business, driving incident troubleshooting and root-cause prevention while designing resilient, secure infrastructure across cloud and container ecosystems. Own SRE practices such as SLIs/SLOs, error budgets, and automation to reduce toil. Build observability and deployment automation (CI/CD, infrastructure-as-code), mentor engineers, and partner across engineering, QA, product, and operations to embed reliability end-to-end.
Location: Guangzhou
Employment Type: Full time
Job Function: Software Engineering
Seniority: Director level

Key Responsibilities

  • •Lead complex troubleshooting and root cause analysis for incidents impacting production to drive rapid resolution and long-term prevention.
  • •Design and enhance scalable, highly available, and secure infrastructure using cloud, containers, and orchestration technologies.
  • •Define and mature SRE practices (SLIs/SLOs, error budgets) and automate operational processes to minimize toil.
  • •Develop monitoring, logging, and alerting systems using observability tools and drive deployment automation and CI/CD improvements.
  • •Guide and mentor engineers, collaborate cross-functionally to embed reliability and security, and lead on-call rotations and blameless postmortems.

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science, Engineering, or related field, or equivalent experience.
  • •At least 10 years of hands-on experience in IT with significant SRE, DevOps, Production Support, or related experience.
  • •Advanced hands-on expertise with containers (Docker, Kubernetes) and cloud platforms (AWS, GCP, Azure).
  • •Deep experience in monitoring, log management, and observability platforms.
  • •Fluent in at least one programming or scripting language (Python, Bash, Go, etc.) and experience implementing SRE at the organizational level.
Experience:SREDevOpsProduction Support
Skills:Problem-solvingAnalyticalCommunicationMentoring
Tech Stack:AWSGCPAzureKubernetesDockerPrometheusGrafanaELKDatadogSplunkTerraformAnsibleHelmCI/CDInfrastructure-as-codeConfiguration managementSLIsSLOsError budgetsObservability

Company Brief

HSBC
Global banking and financial services organisation offering retail, commercial, corporate and investment banking, wealth management, and global markets services across Europe, Asia, the Americas and the Middle East.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: London, United Kingdom
Founded: 1865
Glassdoor
Glassdoor: 3.6
WebsiteLinkedIn