Site Reliability Engr II

Honeywell
Bengaluru
Full timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: []

Design, build, and maintain fault-tolerant, scalable, and highly available distributed systems. Own reliability and availability by defining SLOs/SLIs/SLAs, disaster recovery (RTO/RPO), and capacity planning, while leading incident response and root cause analysis. Improve monitoring and observability with tools like Prometheus, Grafana, Azure Monitor, Dynatrace, and Elastic. Automate operational toil using IaC and enhance CI/CD and cloud infrastructure across Kubernetes/AKS and related services.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Honeywell
Honeywell
4 days ago

Site Reliability Engr II

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Design, build, and maintain fault-tolerant, scalable, and highly available distributed systems. Own reliability and availability by defining SLOs/SLIs/SLAs, disaster recovery (RTO/RPO), and capacity planning, while leading incident response and root cause analysis. Improve monitoring and observability with tools like Prometheus, Grafana, Azure Monitor, Dynatrace, and Elastic. Automate operational toil using IaC and enhance CI/CD and cloud infrastructure across Kubernetes/AKS and related services.
Location: Bengaluru
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure high availability and uptime of production services by defining and managing SLOs, SLIs, and SLAs.
  • •Design and implement disaster recovery strategies (including RTO/RPO) and perform capacity planning/scalability assessments.
  • •Implement monitoring, logging, and alerting; build dashboards and improve observability; detect, investigate, and resolve system issues.
  • •Participate in on-call rotations, lead troubleshooting during service disruptions, and perform root cause analysis to drive corrective actions while reducing MTTD/MTTR.
  • •Develop automation to eliminate operational toil, build self-healing and auto-scaling, maintain IaC, and improve CI/CD pipeline reliability while managing cloud infrastructure and Kubernetes clusters.

Key Requirements

  • •Bachelor’s/Master’s degree in Computer Science, Information Technology, or equivalent practical experience.
  • •5+ years supporting large-scale cloud-native applications.
  • •Experience with incident management and production support.
  • •Understanding of SRE concepts including error budgets, SLI/SLO/SLA, chaos engineering, high availability, and reliability engineering.
  • •Experience contributing to reliability success metrics such as uptime, MTTR/MTTD, deployment success rate, and production incident reduction.
Education:Bachelor's in Computer Science, Information Technology
Languages:US
Tech Stack:AzureAWSGCPKubernetesAKSOpenShiftLinuxTCP/IPDNSLoad BalancingSQLNoSQLTerraformARMBicepCloudFormationAzure DevOpsGitHub ActionsJenkinsGitLab

Company Brief

Honeywell
Global diversified technology and manufacturing company providing aerospace systems, building technologies, performance materials, and safety & productivity solutions for industrial, commercial, and consumer markets.
Industry: Conglomerates & Holding Companies
Company Size: Enterprise (1,001+ employees)
Revenue: USD 25B to 50B
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Charlotte, United States
Founded: 1906
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor