Site Reliability Engineer

Tyk
Canada
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Proactive","Innovation","Continuous improvement"]

Own the reliability of a global API management cloud platform—maintaining production services, identifying reliability issues, and driving incident response. Build operational metrics and dashboards, automate recurring tasks, and expand multi-region/multi-cloud reach. Partner with your squad to improve SLAs/SLOs, conduct post-incident analysis, document SRE processes, and help with cloud penetration testing—while operating Kubernetes, AWS/EKS, Linux, and observability tooling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tyk
Tyk
10 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own the reliability of a global API management cloud platform—maintaining production services, identifying reliability issues, and driving incident response. Build operational metrics and dashboards, automate recurring tasks, and expand multi-region/multi-cloud reach. Partner with your squad to improve SLAs/SLOs, conduct post-incident analysis, document SRE processes, and help with cloud penetration testing—while operating Kubernetes, AWS/EKS, Linux, and observability tooling.
Location: Canada
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Maintain and improve global Tyk Cloud services by defining and operating SLAs/SLOs/SLIs and responding to reliability issues.
  • •Own incident management as the first line for clients, including post-incident analysis and defining response improvements.
  • •Introduce reliability metrics, build dashboards, and participate in the on-call rotation.
  • •Expand the platform’s multi-region and multi-cloud reach with your squad while automating common operational tasks.
  • •Document SRE processes and policies, recommend operational efficiency improvements, and assist in cloud penetration testing.

Pay and Benefits

Perks:Paid LeaveEquityParental Leave

Key Requirements

  • •Strong collaboration skills and comfort working in an operational/on-call environment.
  • •Experience launching and operating production-scale Kubernetes clusters.
  • •Designing and operating infrastructure on AWS (EKS) and other providers.
  • •Operational experience with MongoDB and Redis cluster administration.
  • •Experience with monitoring/observability (Prometheus, Grafana) and logging collection/analysis.
Skills:CollaborationProactiveInnovationContinuous improvement
Certifications:CKACKADCKS
Tech Stack:KubernetesContainersAWSEKSLinuxTerraformIaCHelmGoMongoDBRedisPrometheusGrafanaThanosSubnetsRoutingPeeringLoad balancingNATDNS

Company Brief

Tyk
Tyk provides an open-source API gateway and API management platform (self-managed, cloud, hybrid) offering gateway, analytics, developer portal and dashboard functionality to help enterprises secure, monitor and manage APIs at scale.
Industry: API Platforms
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Series B
Headquarters: London, United Kingdom
Founded: 2014
Glassdoor
Glassdoor: 4.1
WebsiteLinkedInGlassdoor