Senior Site Reliability Engineer

2K
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Leadership","Ownership","Systemic thinking"]

Own and advance production reliability for 2K’s live, multi-cloud game services. Design, build, and operate scalable infrastructure across AWS, GCP, and on-prem using Terraform, Pulumi, and GitOps (ArgoCD/Flux). Run Kubernetes platforms end-to-end (EKS/GKE) with service mesh and progressive delivery. Build observability with Prometheus, Grafana, Datadog, and OpenTelemetry, lead SLI/SLO and chaos engineering, and drive systemic incident and post-mortem improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
2K
2K
2 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 38 minutes agoStatus: Live

Job Summary

Own and advance production reliability for 2K’s live, multi-cloud game services. Design, build, and operate scalable infrastructure across AWS, GCP, and on-prem using Terraform, Pulumi, and GitOps (ArgoCD/Flux). Run Kubernetes platforms end-to-end (EKS/GKE) with service mesh and progressive delivery. Build observability with Prometheus, Grafana, Datadog, and OpenTelemetry, lead SLI/SLO and chaos engineering, and drive systemic incident and post-mortem improvements.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate scalable multi-cloud and hybrid infrastructure, including Kubernetes platform ownership (cluster lifecycle, multi-tenancy, networking, autoscaling).
  • •Implement progressive delivery and reliability patterns across game service deployments (blue/green, canary).
  • •Build and run the observability stack (Prometheus, Grafana, Datadog, OpenTelemetry) and define SLI/SLO/error budgets with signal-based alerting.
  • •Lead chaos engineering and drive incident response and post-mortems focused on systemic fixes and follow-through.
  • •Reduce toil via automation (self-service provisioning, automated remediation, intelligent scaling) and harden CI/CD pipelines and platform-layer security (secrets management, policy-as-code).

Pay and Benefits

Perks:Annual BonusProvident FundHealth InsuranceLife AssuranceAccident InsuranceChildcareGym MembershipLearning Platforms

Key Requirements

  • •5+ years in SRE, Platform Engineering, or equivalent infrastructure work at production scale.
  • •Deep Kubernetes experience in cloud environments (EKS or GKE preferred), including networking, storage, and multi-cluster patterns.
  • •Strong IaC skills with Terraform and/or Pulumi, with hands-on Helm, Terragrunt, and GitOps tooling (ArgoCD or GitHub Actions).
  • •Observability experience with Datadog plus Prometheus/Grafana and OpenTelemetry, including SLI/SLO/error budget operationalization.
  • •Production-quality coding in Go, Python, or TypeScript, plus Linux internals and core network fundamentals for system-level debugging.
Experience:5+ yearsLive-service gamesConsumer internetProduction scale
Skills:LeadershipOwnershipSystemic thinking
Certifications:AWS Solutions ArchitectGCP Professional Cloud ArchitectCKACKS
Languages:English
Tech Stack:AWSGCPTerraformPulumiGitOpsArgoCDFluxKubernetesEKSGKEIstioCiliumPrometheusGrafanaDatadogOpenTelemetrySLISLOBlue/greenCanary

Company Brief

2K
2K is a video game publisher and developer known for sports, action, and strategy franchises such as NBA 2K, WWE 2K, Borderlands, and Sid Meier’s Civilization. It operates as a major label within Take-Two Interactive.
Industry: Gaming
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Novato, United States
Founded: 2005
WebsiteLinkedIn