Software Engineer II - Platform & Infrastructure

Abnormal AI
Bengaluru
Workplace: HybridFull timeFunction: Software EngineeringExperience: 4+ yearsSkills: ["Async-first communication","Mentoring","Code review","Problem-solving","Technical documentation"]

Build and evolve the Platform & Infrastructure responsible for Abnormal’s observability and data infrastructure. Own the monitoring, metrics, alerting, and incident response stack spanning Prometheus/Chronosphere/Grafana and the PagerDuty alerting pipeline, while also driving reliability, SLAs/SLOs, and cost-efficient operations across US, EU, and GovCloud. Design end-to-end platform features, improve operational automation, mentor peers, and create tooling that helps other engineers ship faster.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Abnormal AI
Abnormal AI
3 months ago

Software Engineer II - Platform & Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and evolve the Platform & Infrastructure responsible for Abnormal’s observability and data infrastructure. Own the monitoring, metrics, alerting, and incident response stack spanning Prometheus/Chronosphere/Grafana and the PagerDuty alerting pipeline, while also driving reliability, SLAs/SLOs, and cost-efficient operations across US, EU, and GovCloud. Design end-to-end platform features, improve operational automation, mentor peers, and create tooling that helps other engineers ship faster.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own and evolve the observability platform, including monitoring, metrics, and alerting infrastructure across Prometheus/Chronosphere/Grafana and the PagerDuty alerting pipeline.
  • •Design, develop, and deliver platform features end-to-end from technical design docs through production rollout and post-launch monitoring across US, EU, and GovCloud.
  • •Take ownership of 1–3 key services within Observability or Data Infra, driving reliability, performance, and evolution.
  • •Participate in on-call rotations to triage, diagnose, and resolve production issues independently while building deep operational knowledge.
  • •Improve resilience and reduce operational toil by automating runbooks, refining SLAs/SLOs, and strengthening testing and observability instrumentation.

Key Requirements

  • •4+ years of hands-on backend engineering experience designing, building, and operating production-grade distributed systems.
  • •Strong proficiency in Python and working proficiency in Golang for infrastructure components, metric pipelines, and platform services.
  • •Experience building and owning services end-to-end, including technical design, production deployment, monitoring, and iteration.
  • •Demonstrated incident response capability from on-call experience, with strong testing discipline (unit and integration tests).
  • •Solid understanding of monitoring/alerting/observability principles and fault tolerance patterns (retries, circuit breakers, graceful degradation, backpressure).
Experience:4+ years
Skills:Async-first communicationMentoringCode reviewProblem-solvingTechnical documentation
Languages:English
Tech Stack:PythonGolangPrometheusChronosphereGrafanaPagerDutyPromQLAirflowSparkAWSEC2ECSEKSS3RDSIAMCloudWatchLambdaSQSSNS

Company Brief

Abnormal AI
Provides AI-powered email security and human behavior analysis to detect phishing, business email compromise, and account takeover attacks. The platform helps organizations protect employees and communications across cloud email and collaboration tools.
Industry: Cybersecurity
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: San Francisco, United States
Founded: 2018
WebsiteLinkedIn