Lead Software Engineer - Site Reliability

Freshworks
Hyderabad
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 7-12 yearsEducation: bachelorsSkills: ["Problem-solving","Diagnostic skills","Communication","Automation mindset","Works in high-pressure environments"]

Design and lead reliability engineering for production systems, focusing on resilience, automation, and observability. Define SLIs/SLOs and error budgets, build monitoring/alerting and remediation pipelines, and drive performance and availability improvements. Partner with engineering, platform, and product teams to shift reliability left, lead incident response and blameless postmortems, and contribute to infrastructure architecture, automation, and reliability roadmaps at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Freshworks
Freshworks
1 day ago

Lead Software Engineer - Site Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live
Reposted: similar role first listed 8 months ago

Job Summary

Design and lead reliability engineering for production systems, focusing on resilience, automation, and observability. Define SLIs/SLOs and error budgets, build monitoring/alerting and remediation pipelines, and drive performance and availability improvements. Partner with engineering, platform, and product teams to shift reliability left, lead incident response and blameless postmortems, and contribute to infrastructure architecture, automation, and reliability roadmaps at scale.
Location: Hyderabad
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement tools to improve availability, latency, scalability, and overall system health.
  • •Define SLIs/SLOs, manage error budgets, and drive performance engineering efforts.
  • •Build and maintain automated monitoring, alerting, and remediation pipelines.
  • •Lead incident response, root cause analysis, and blameless postmortems to improve reliability.
  • •Champion observability and contribute to infrastructure architecture, automation, and reliability roadmaps.

Key Requirements

  • •7–12 years of experience in SRE, DevOps, or Production Engineering roles.
  • •Strong coding proficiency to develop clear, efficient, well-structured code.
  • •In-depth Linux expertise for system administration and advanced troubleshooting.
  • •Practical experience with Docker and Kubernetes, plus CI/CD pipeline design, implementation, and maintenance.
  • •Experience across high availability/scalability, disaster recovery, observability, and infrastructure automation (IaC) with security best practices.
Experience:7-12 yearsSREDevOpsProduction EngineeringDistributed systems
Education:Bachelor's in Computer Science, Engineering
Skills:Problem-solvingDiagnostic skillsCommunicationAutomation mindsetWorks in high-pressure environments
Languages:English
Tech Stack:LinuxDockerKubernetesCI/CDContinuous IntegrationContinuous DeliveryInfrastructure as Code (IaC)SLIsSLOsError budgetsMonitoringLoggingTracingObservabilityIncident responseRoot cause analysisBlameless postmortemsDisaster recoveryHigh availabilityDistributed systems

Company Brief

Freshworks
Provides cloud-based customer engagement, CRM and ITSM software (Freshdesk, Freshservice, Freshsales) that helps businesses manage customer and employee experiences through SaaS products and AI-driven tools.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 500M to 1B
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Mateo, United States
Founded: 2010
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor