Site Reliability Engineer (SRE)

TTEC Digital
Hyderabad
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Self-starter","Grit","Calm in incident","Relentless follow-through","Teamwork"]

Own production reliability for a real-time platform where uptime and latency are the product, spanning voice, desktop, intelligence, and AI. You’ll set SLOs and error budgets per tenant/service, lead incident response with blameless postmortems, and drive production scaling and capacity. Build deep observability (p50/p95/p99), participate in on-call with DevOps, and help ship weekly in a fast startup environment through deploy-safety practices and chaos engineering.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TTEC Digital
TTEC Digital
1 month ago

Site Reliability Engineer (SRE)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Own production reliability for a real-time platform where uptime and latency are the product, spanning voice, desktop, intelligence, and AI. You’ll set SLOs and error budgets per tenant/service, lead incident response with blameless postmortems, and drive production scaling and capacity. Build deep observability (p50/p95/p99), participate in on-call with DevOps, and help ship weekly in a fast startup environment through deploy-safety practices and chaos engineering.
Location: Hyderabad
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own production reliability for a real-time platform where uptime and latency are the product.
  • •Set and manage SLOs and error budgets per tenant/service.
  • •Run incident response and lead blameless postmortems.
  • •Drive production scaling and capacity planning.
  • •Partner on observability and tenancy isolation, including on-call rotation with DevOps.

Key Requirements

  • •8+ years operating production systems at scale, owning SLOs, error budgets, and incident command.
  • •Strong Go or Python skills to automate reliability with runbooks and automated remediation.
  • •Deep experience with event-driven and real-time reliability, including debugging failure physics in production.
  • •Strong monitoring and uptime mindset with metrics, logs, and traces wired to alerting and signal detection.
  • •Good networking understanding (TCP/UDP, TLS, WebSocket, DNS, load balancing) with RTP/SIP as a plus.
Experience:Real-time systemsEvent-driven systemsProduction reliability
Skills:Self-starterGritCalm in incidentRelentless follow-throughTeamwork
Tech Stack:GoPythonNATS-class busesWebSocketStreaming pipelinesTCPUDPTLSDNSLoad balancingRTPSIPGCPMulti-cloudObservabilitySLOsError budgetsIncident responseBlameless postmortemsP50

Company Brief

TTEC Digital
Provides digital customer experience consulting, technology and managed services, delivering CX platforms, automation, analytics, and design to help enterprises transform customer engagement across channels.
Industry: Consulting
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Englewood, United States
WebsiteLinkedIn