Staff Software Engineer, Reliability - Command|Alert

CommandLink
Argentina, Brazil, Chile, Colombia, Costa Rica, El Salvador, India, Mexico, Philippines
Workplace: RemoteFull timeFunction: Software EngineeringSkills: ["Incident response","Blameless post-mortems","Systems thinking","Mentoring","Stakeholder management"]

Own end-to-end reliability for the alerting pipeline that converts security, monitoring, and customer telemetry into trustworthy alerts. Define SLOs/SLIs and drive improvements using error budgets, run blameless incident response, and build observability/automation to keep operational toil low at Fortune 1000 scale. Architect rule evaluation and correlation logic across OpenSearch, Kafka, and security telemetry, mentoring engineers and shaping technical direction.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
CommandLink
CommandLink
2 days ago

Staff Software Engineer, Reliability - Command|Alert

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own end-to-end reliability for the alerting pipeline that converts security, monitoring, and customer telemetry into trustworthy alerts. Define SLOs/SLIs and drive improvements using error budgets, run blameless incident response, and build observability/automation to keep operational toil low at Fortune 1000 scale. Architect rule evaluation and correlation logic across OpenSearch, Kafka, and security telemetry, mentoring engineers and shaping technical direction.
Location: Argentina, Brazil, Chile, Colombia, Costa Rica, El Salvador, India, Mexico, Philippines
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own reliability of the alerting pipeline end to end, including idempotency guarantees and soak-tested behavior from OpenSearch alert evaluation through Kafka delivery and downstream notifications.
  • •Define SLOs/SLIs for availability, latency, and delivery guarantees, using error budgets to balance reliability investment against new capability.
  • •Lead blameless incident response for the pipeline’s most critical production issues and drive systemic fixes and automation from post-mortems.
  • •Build observability and automation to run the pipeline at Fortune 1000 scale with low operational toil, and own capacity planning as ingestion and customer counts grow.
  • •Set the architecture for evaluating rule-based thresholds, ML anomaly scores, and correlation logic to turn telemetry into network/system topologies for investigation and remediation.

Key Requirements

  • •Background operating as a Site Reliability Engineer, DevOps engineer, or similar production-ownership role with fluency in SLOs, SLIs, error budgets, and blameless incident response.
  • •Experience designing, building, or operating high-reliability alerting or notification systems in production, including rule-based and ML-based detection at scale.
  • •Strong Kafka experience and a track record building systems with webhook reliability, idempotency, and delivery guarantees under load.
  • •Experience building observability and automation so a high-volume production system runs with low operational toil.
  • •Command of telemetry and protocol data (e.g., security tooling output, syslog, OpenTelemetry, NetFlow/sFlow, SNMP, ICMP, firewall logs) to produce usable network and system topologies.
Experience:SREDevOpsHigh-reliability systemsAlerting/notification systemsMulti-cloudCloud-nativeTelemetrySecurity monitoring
Skills:Incident responseBlameless post-mortemsSystems thinkingMentoringStakeholder management
Tech Stack:OpenSearchKafkaOpenTelemetrySyslogNetFlow/sFlowSNMPICMPFirewall logsGoPythonContainer orchestrationMulti-cloudRule-based thresholdsML anomaly scoresCorrelation logicTLSSecurity tooling outputL2-L4 network protocols

Company Brief

CommandLink
Provides secure cloud, managed IT, contact center, and cybersecurity solutions tailored for enterprises and government clients, delivering mission-critical communications, infrastructure, and support services to enable secure operations and customer engagement.
Industry: Telecommunications
Website