Site Reliability Engineer (US - Pacific time)

PostHog
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Proactive","Problem-solving","Collaboration","On-call"]

SRE role focused on turning a fast-growing, stateful system into a predictable, well-automated platform. You’ll own production Kubernetes (EKS), multi-account AWS, Terraform/Terragrunt IaC, and tooling for deployments, backups, and incident response, aiming to reduce operational load and enable scalable, reliable services across regions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
PostHog
PostHog
4 months ago

Site Reliability Engineer (US - Pacific time)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 minutes agoStatus: Live

Job Summary

SRE role focused on turning a fast-growing, stateful system into a predictable, well-automated platform. You’ll own production Kubernetes (EKS), multi-account AWS, Terraform/Terragrunt IaC, and tooling for deployments, backups, and incident response, aiming to reduce operational load and enable scalable, reliable services across regions.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and manage EKS clusters across environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
  • •Develop and evolve a multi-AWS account organization, provisioning, networking, access control, and cross-account connectivity
  • •Maintain the Terraform/Terragrunt IaC platform—modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
  • •Improve operational tooling around deploys, schema changes, backups, restores, and incident response
  • •Reduce operational load by identifying repeat pain points and eliminating them through code and self-healing automation

Key Requirements

  • •Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
  • •Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
  • •Experience automating infrastructure using Terraform or Terragrunt at scale, including module design and state management
  • •Solid understanding of Linux systems (disk, memory, networking, failure modes)
  • •Experience supporting stateful systems (databases, queues, storage systems, etc.)
Experience:AWSKubernetesEKSTerraformTerragruntLinuxStateful systemsIaCGitOps
Skills:OwnershipProactiveProblem-solvingCollaborationOn-call
Languages:English
Tech Stack:KubernetesEKSAWSTerraformTerragruntLinuxArgoCDGitOpsKarpenterCiliumGitHub Actions

Company Brief

PostHog
PostHog provides an open-source product analytics and developer platform that combines analytics, session replay, feature flags, A/B testing, and data warehousing to help engineering and product teams build and optimize products.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2020
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor