Customer Reliability Engineer, Infrastructure

Astronomer
San Francisco, Austin, New York, Boston
Workplace: RemoteFull timeUSD 125,000 - 130,000 annuallyFunction: Data Analytics & Business IntelligenceExperience: 5+ yearsSkills: ["Communication","Troubleshooting","Customer-facing support","Incident triage"]

Own the reliability of Astronomer’s managed Airflow service by operating, monitoring, and maintaining underlying cloud infrastructure and Kubernetes clusters. Respond to customer- and monitoring-raised incidents, drive permanent fixes, and improve observability with better monitoring/alerting and automation. Work directly with customers to triage issues, meet SLAs, provide “white glove” guidance, and feed product teams with recurring needs and pain points across multi-cloud environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Astronomer
Astronomer
22 hours ago

Customer Reliability Engineer, Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Own the reliability of Astronomer’s managed Airflow service by operating, monitoring, and maintaining underlying cloud infrastructure and Kubernetes clusters. Respond to customer- and monitoring-raised incidents, drive permanent fixes, and improve observability with better monitoring/alerting and automation. Work directly with customers to triage issues, meet SLAs, provide “white glove” guidance, and feed product teams with recurring needs and pain points across multi-cloud environments.
Location: San Francisco, Austin, New York, Boston
Workplace: Remote
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Mid level

Key Responsibilities

  • •Operate, monitor, and maintain the managed Airflow platform to ensure availability, predictability, and reliable operations.
  • •Respond to incidents raised by customers or monitoring systems and drive steps to permanently resolve or monitor problems.
  • •Provide customer-facing solutions: troubleshoot environments, triage issues with customers, and meet SLAs with white-glove guidance.
  • •Build and maintain monitoring/alerting systems and automation for efficient daily operational tasks.
  • •Collaborate with product development by feeding back customer needs and pain points, and help direct product architecture where possible.

Pay and Benefits

Salary: USD 125,000 - 130,000 annually
Equity and Bonus:Equity

Key Requirements

  • •5+ years of experience operating large, complex cloud infrastructures at scale.
  • •3+ years of experience with Kubernetes.
  • •Experience managing a production distributed system with AWS, GCP, and/or Azure.
  • •Strong Linux experience and knowledge of operating/monitoring distributed systems.
  • •DevOps or CI/CD experience, including Python scripting and strong troubleshooting and customer issue handling.
Experience:5+ yearsCloud infrastructureDistributed systemsMulti-cloudKubernetesDevOpsCI/CDData infrastructure
Skills:CommunicationTroubleshootingCustomer-facing supportIncident triage
Tech Stack:Apache AirflowKubernetesLinuxAWSGCPAzurePythonCI/CDIaCObservabilityMonitoringAlertingAutomation

Company Brief

Astronomer
Develops Astro, a unified DataOps platform powered by Apache Airflow to build, run, and monitor data pipelines at scale, helping enterprises operationalize analytics and production AI.
Industry: Data Infrastructure
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series D
Headquarters: New York, United States
Founded: 2018
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor