Senior Staff Engineer - Site Reliability

Freshworks
Chennai
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10-15 yearsSkills: ["Stakeholder communication","Requirements gathering","Documentation","Problem-solving","Cross-functional collaboration"]

Build and evolve Freshworks’ SRE practice to improve availability, latency, and efficiency across Products & Platforms. Design scalable, cloud-native architectures, implement self-healing and auto-scaling, and manage performance for mission-critical services. Own automation and orchestration strategies, drive blameless postmortems and reliability goals, and optimize cloud costs. Work with cross-functional teams to execute technical roadmaps using SLOs, monitoring, telemetry, and infrastructure-as-code.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Freshworks
Freshworks
5 days ago

Senior Staff Engineer - Site Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and evolve Freshworks’ SRE practice to improve availability, latency, and efficiency across Products & Platforms. Design scalable, cloud-native architectures, implement self-healing and auto-scaling, and manage performance for mission-critical services. Own automation and orchestration strategies, drive blameless postmortems and reliability goals, and optimize cloud costs. Work with cross-functional teams to execute technical roadmaps using SLOs, monitoring, telemetry, and infrastructure-as-code.
Location: Chennai
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, write, and deliver software to improve availability, latency, and efficiency of Freshworks products and platforms.
  • •Develop scalable, cloud-native architectures supporting business growth.
  • •Design and implement self-healing and auto-scaling mechanisms to prevent recurring problems.
  • •Manage availability, latency, and performance for mission-critical services and build automation to prevent problem recurrence.
  • •Drive reliability roadmaps, automation/orchestration strategies, and cloud cost optimization, including blameless postmortems for large-scale incidents.

Key Requirements

  • •10–15 years of experience in SRE, handling performance, architecture, and application design.
  • •Strong understanding of cloud computing, networking, Linux systems administration, containerization (Docker, Kubernetes), and infrastructure as code (Terraform, Ansible).
  • •Experience with SRE principles including SLOs, SLIs, SLAs, and error budgets.
  • •Experience managing incidents and retrospectives, including blameless postmortems for large-scale incidents.
  • •Hands-on monitoring, logging & telemetry experience using tools like New Relic, Splunk, ELK, Nagios, SolarWinds, Prometheus, AWS CloudWatch, Datadog, and OpenTelemetry.
Experience:10-15 yearsSRECloud computingDistributed systemsPerformance engineeringInfrastructure as code
Skills:Stakeholder communicationRequirements gatheringDocumentationProblem-solvingCross-functional collaboration
Languages:English (US)
Tech Stack:AWSLinuxDockerKubernetesTerraformAnsibleNew RelicSplunkELKNagiosSolarWindsPrometheusAWS CloudwatchDatadogOpentelemetry

Company Brief

Freshworks
Provides cloud-based customer engagement, CRM and ITSM software (Freshdesk, Freshservice, Freshsales) that helps businesses manage customer and employee experiences through SaaS products and AI-driven tools.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 500M to 1B
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Mateo, United States
Founded: 2010
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor