Senior Site Reliability Engineer

2K
Burnaby
Workplace: HybridFull timeUSD 93,700 - 138,700 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Leadership","Incident response","Systemic problem-solving","Collaboration","Ownership"]

Own the 2K SRE platform that keeps player-facing systems running across AWS, GCP, and on-prem data centers. Design and operate scalable multi-cloud and hybrid infrastructure, including Kubernetes (EKS/GKE) and networking, and drive progressive delivery. Build and run observability (Prometheus/Grafana/Datadog), SLI/SLOs, alerting, and incident response with systemic post-mortems. Lead automation, security, and developer experience through CI/CD hardening and policy-as-code.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
2K
2K
3 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own the 2K SRE platform that keeps player-facing systems running across AWS, GCP, and on-prem data centers. Design and operate scalable multi-cloud and hybrid infrastructure, including Kubernetes (EKS/GKE) and networking, and drive progressive delivery. Build and run observability (Prometheus/Grafana/Datadog), SLI/SLOs, alerting, and incident response with systemic post-mortems. Lead automation, security, and developer experience through CI/CD hardening and policy-as-code.
Location: Burnaby
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate scalable multi-cloud and hybrid infrastructure across AWS, GCP, and on-premises data centers.
  • •Own Kubernetes platform lifecycle end-to-end, including multi-tenancy, networking, and autoscaling, and drive progressive delivery for game service deployments.
  • •Build and run the observability stack (Prometheus/Grafana/Datadog), define SLI/SLO/error budget policies, and implement alerting that reduces noise.
  • •Lead chaos engineering exercises, and drive incident response and post-mortems focused on systemic fixes and follow-through.
  • •Promote reliability through runbooks and reliability reviews, author engineering RFCs, and lead automation/security initiatives across CI/CD pipelines and secrets/policy-as-code.

Pay and Benefits

Salary: USD 93,700 - 138,700 annually
Equity and Bonus:Equity

Key Requirements

  • •5+ years in SRE, Platform Engineering, or equivalent infrastructure work at production scale.
  • •Deep Kubernetes experience in cloud environments (EKS or GKE preferred), including networking, storage, and multi-cluster patterns.
  • •Strong IaC proficiency with Terraform and/or Pulumi, with hands-on experience using Helm, Terragrunt, and GitOps tooling (ArgoCD or GitHub Actions).
  • •Experience with modern and legacy infrastructure including AWS, GCP, VMware, and bare metal servers.
  • •Production-quality code in Go, Python, or TypeScript plus observability experience (Datadog, Prometheus + Grafana, OpenTelemetry) and SLI/SLO/error budget fluency.
Experience:SREPlatform engineeringKubernetesAWSGCPLive-serviceConsumer internet
Skills:LeadershipIncident responseSystemic problem-solvingCollaborationOwnership
Certifications:AWS Solutions ArchitectGCP Professional Cloud ArchitectCKACKS
Tech Stack:AWSGCPTerraformPulumiGitOpsArgoCDFluxKubernetesEKSGKEIstioCiliumBlue/greenCanaryPrometheusGrafanaDatadogOpenTelemetrySLISLO

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

2K
2K is a video game publisher and developer known for sports, action, and strategy franchises such as NBA 2K, WWE 2K, Borderlands, and Sid Meier’s Civilization. It operates as a major label within Take-Two Interactive.
Industry: Gaming
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Novato, United States
Founded: 2005
WebsiteLinkedIn