Major Incident Manager

Sonar Source
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Change management","Infrastructure management","System support","Configuration management","Incident response"]

Drive major incident operations and engineering improvements by using and building automation to monitor and observe production infrastructure across on-prem and cloud. Own system health monitoring, alert triaging, SLO/SLA-style error budget prioritization, and incident response through engineering actions from post-mortems. Strengthen security and observability through DevSecOps pipeline maintenance, infrastructure/policy as code, and continuous toil-elimination to improve reliability and MTTR.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Sonar Source
Sonar Source
3 days ago

Major Incident Manager

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Drive major incident operations and engineering improvements by using and building automation to monitor and observe production infrastructure across on-prem and cloud. Own system health monitoring, alert triaging, SLO/SLA-style error budget prioritization, and incident response through engineering actions from post-mortems. Strengthen security and observability through DevSecOps pipeline maintenance, infrastructure/policy as code, and continuous toil-elimination to improve reliability and MTTR.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Monitor and observe production infrastructure services on premises and in the cloud using automation tools.
  • •Triaging high-severity alerts, managing error budget burn rate, and prioritizing work using SLO-linked dashboards.
  • •Develop, review, and test infrastructure and security automation using IaC and policy-as-code practices to prevent configuration drift.
  • •Eliminate operational toil by automating repetitive security patching, compliance checks, certificate rotations, and infrastructure maintenance.
  • •Maintain DevSecOps security tooling in CI/CD pipelines and participate in on-call incident response, translating post-mortems into preventative code and automated runbook actions.

Key Requirements

  • •Deep IaC experience provisioning and managing complex infrastructure using Terraform or CloudFormation (AWS) and/or configuration management tools like Ansible or Puppet.
  • •Hands-on cloud/platform experience with a major cloud provider (AWS, GCP, or Azure) or managing large-scale internal/private cloud infrastructure.
  • •Experience defining, measuring, and reporting Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for critical services.
  • •Proven observability experience with tools such as Prometheus/Grafana, ELK/EFK stacks, and/or vendor solutions like Datadog or Splunk.
  • •Incident management experience, including contributing to post-incident reviews and prioritizing engineering work from error budget management.
Skills:Change managementInfrastructure managementSystem supportConfiguration managementIncident response
Tech Stack:PythonGoTerraformCloudFormationAWSGCPAzureAnsiblePuppetPrometheusGrafanaELKEFKDatadogSplunkHashiCorp Vault

Company Brief

Sonar Source
Builds static code analysis and continuous inspection tools (SonarQube, SonarCloud, SonarLint) that identify bugs, vulnerabilities, and code smells across multiple languages to help teams improve code quality and maintainability.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Geneva, Switzerland
Founded: 2008
WebsiteLinkedIn