Senior Customer Reliability Engineer, Infrastructure - Hyderabad, India

Astronomer
Hyderabad
Workplace: HybridFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsSkills: ["Communication","Troubleshooting"]

Own reliability for a managed Airflow service by operating, monitoring, and improving cloud infrastructure and Kubernetes clusters. You’ll troubleshoot customer-raised and monitoring incidents, build monitoring/alerting and automation for daily operations, and help maintain 24x7 coverage through paid on-call. As owners of the observability platform, you’ll provide white-glove guidance, meet SLAs, and feed customer feedback into product and architecture improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Astronomer
Astronomer
4 days ago

Senior Customer Reliability Engineer, Infrastructure - Hyderabad, India

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own reliability for a managed Airflow service by operating, monitoring, and improving cloud infrastructure and Kubernetes clusters. You’ll troubleshoot customer-raised and monitoring incidents, build monitoring/alerting and automation for daily operations, and help maintain 24x7 coverage through paid on-call. As owners of the observability platform, you’ll provide white-glove guidance, meet SLAs, and feed customer feedback into product and architecture improvements.
Location: Hyderabad
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Operate, monitor, and maintain the managed Airflow platform to ensure availability, predictability, and reliable operations.
  • •Respond to incidents from customers or monitoring, triage issues, and drive steps to permanently resolve or properly monitor them.
  • •Build and maintain monitoring and alerting systems, plus automation to handle daily operational tasks efficiently.
  • •Provide white-glove, customer-facing guidance to prioritize and solve issues, meet SLAs, and contribute to production success.
  • •Contribute to architecture and enhance customer documentation while participating in distributed-team on-call coverage.

Key Requirements

  • •5+ years of experience operating large, complex cloud infrastructures at scale.
  • •3+ years of experience with Kubernetes.
  • •Experience managing a production distributed system with at least one major cloud provider (AWS, GCP, or Azure).
  • •Strong Linux and network experience, plus knowledge of operating and monitoring distributed systems.
  • •Experience with observability tools and handling customer issues (internal and external).
Experience:5+ yearsCloud infrastructure
Skills:CommunicationTroubleshooting
Tech Stack:KubernetesAWSGCPAzureLinuxPythonCI/CDDevOpsObservabilityMonitoringAlertingAutomationAirflowBig data orchestrationIaC

Company Brief

Astronomer
Develops Astro, a unified DataOps platform powered by Apache Airflow to build, run, and monitor data pipelines at scale, helping enterprises operationalize analytics and production AI.
Industry: Data Infrastructure
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series D
Headquarters: New York, United States
Founded: 2018
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor