Site Reliability Engineer (Contract, Rotation-Based)

Invisible Technologies
Estonia, India, South Africa
Workplace: RemoteContractFunction: DevOps, Cloud & InfrastructureSkills: ["Incident response","Triage","Troubleshooting","Clear communication","Calm decision-making"]

Serve as a first responder in a 24/7 incident response rotation for a production platform used by a key client. Triage and stabilize infrastructure-level issues within defined SLAs by diagnosing failures from system logs (Kubernetes, RabbitMQ, Postgres). Escalate application/business-logic issues with clear context, and for infra problems, identify and submit fixes. Communicate status to stakeholders and participate in weekend coverage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Invisible Technologies
Invisible Technologies
1 week ago

Site Reliability Engineer (Contract, Rotation-Based)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Serve as a first responder in a 24/7 incident response rotation for a production platform used by a key client. Triage and stabilize infrastructure-level issues within defined SLAs by diagnosing failures from system logs (Kubernetes, RabbitMQ, Postgres). Escalate application/business-logic issues with clear context, and for infra problems, identify and submit fixes. Communicate status to stakeholders and participate in weekend coverage.
Location: Estonia, India, South Africa
Workplace: Remote
Employment Type: Contract
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Serve as first responder for production incidents, triaging and stabilizing at the infrastructure level within defined SLAs (P1/P2).
  • •Diagnose issues primarily from system logs to place failures into the right component before escalating.
  • •Differentiate infrastructure-level failures from application/business-logic failures; submit changes for infra fixes or escalate with clear context.
  • •Participate in an on-call rotation, including off-hours coverage.
  • •Communicate incident status to stakeholders during active incidents and hand off cleanly to the resolution owner.

Key Requirements

  • •Hands-on experience with Kubernetes, RabbitMQ, and PostgreSQL in an enterprise setting, ideally in Financial Services.
  • •Strong working knowledge of Azure; familiarity with GCP or AWS is a plus.
  • •Ability to diagnose unfamiliar systems primarily from logs rather than source code.
  • •Experience troubleshooting production systems under time pressure with good judgment on severity and escalation.
  • •Clear, calm communication during live incidents.
Skills:Incident responseTriageTroubleshootingClear communicationCalm decision-making
Tech Stack:KubernetesRabbitMQPostgresPostgreSQLAzureGCPAWSSystem logsP1P2

Company Brief

Invisible Technologies
Invisible Technologies provides an enterprise AI platform that structures messy data, builds digital workflows, deploys agentic solutions, and integrates human expertise to operationalize AI for customers across industries. Founded in 2015, it has rapidly grown and serves large enterprise clients.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Revenue: USD 100M to 250M
Growth: Scaleup
Funding: Series A
Headquarters: New York, United States
Founded: 2015
WebsiteLinkedIn