Senior Site Reliability Engineer

Carousell Group
Kuala Lumpur
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Troubleshooting","Independent execution","Incident response","Operational improvement","Blameless postmortems"]

Own reliability by operating SLI/SLO/error-budget targets, running end-to-end incident response, and driving blameless postmortems with tracked follow-through. Improve observability by extending instrumentation and raising dashboard/alert quality. Build safe release processes with canary and automated rollback, reduce toil through automation, and manage infrastructure as code. Lead production readiness reviews, performance/capacity planning, and infrastructure security. Operate the MCP gateway and apply AI-assisted tooling to operational workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Carousell Group
Carousell Group
2 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own reliability by operating SLI/SLO/error-budget targets, running end-to-end incident response, and driving blameless postmortems with tracked follow-through. Improve observability by extending instrumentation and raising dashboard/alert quality. Build safe release processes with canary and automated rollback, reduce toil through automation, and manage infrastructure as code. Lead production readiness reviews, performance/capacity planning, and infrastructure security. Operate the MCP gateway and apply AI-assisted tooling to operational workflows.
Location: Kuala Lumpur
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate against SLI/SLOs and error budgets, hold services to targets, and arbitrate between reliability work and feature velocity.
  • •Own the incident lifecycle end to end, including detection, response, mitigation, blameless postmortems, and tracked follow-through across the stack.
  • •Extend and improve observability by instrumenting new services, closing coverage gaps in metrics/logs/traces, and improving dashboard/alert quality.
  • •Design and maintain safe release processes (canary, progressive rollout, automated rollback across dev/staging/production).
  • •Eliminate toil through automation; manage infrastructure as code provisioning/configuration/policies and related documentation.

Key Requirements

  • •5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.
  • •Production experience with Kubernetes deployment, upgrades, troubleshooting and maintenance.
  • •Strong Linux fundamentals and performance tuning, plus Docker and container networking.
  • •Hands-on Google Cloud Platform experience with infrastructure as code and declarative provisioning (Terraform or equivalent).
  • •Experience with secrets and credential management using HashiCorp Vault (or equivalent) and least-privilege access.
Experience:5+ yearsSREDevOpsInfrastructure engineeringProduction systemsKubernetesCloud infrastructureObservability
Skills:TroubleshootingIndependent executionIncident responseOperational improvementBlameless postmortems
Languages:English
Tech Stack:KubernetesLinuxRHELCentOSDebianUbuntuDockerContainer networkingGoogle Cloud PlatformTerraformHashiCorp VaultCI/CDGitHub ActionsBashGoPythonPrometheusGrafanaGoogle Cloud OperationsTracing

Company Brief

Carousell Group
Carousell Group operates a classifieds and C2C marketplace across Southeast Asia, enabling users to buy and sell pre-loved items and providing associated services like listings, payments and buyer protection.
Industry: Online Marketplaces
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Headquarters: Singapore, Singapore
Founded: 2012
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor