Site Reliability Engineer (US - Central/Eastern time)

PostHog
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Deep ownership","Problem-solving","Debugging","Collaboration","Automation mindset","Reliability focus"]

Own reliability for a fast-growing, stateful platform across a multi-region, multi-account AWS environment running many services on Kubernetes. Operate and automate EKS clusters with Karpenter, Cilium, and ArgoCD/GitOps, evolve cross-account infrastructure and access, and maintain Terraform/Terragrunt IaC with safe CI-based workflows. Improve deploy, backup/restore, schema-change, and incident-response tooling while reducing operational load and cloud spend.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
PostHog
PostHog
3 weeks ago

Site Reliability Engineer (US - Central/Eastern time)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own reliability for a fast-growing, stateful platform across a multi-region, multi-account AWS environment running many services on Kubernetes. Operate and automate EKS clusters with Karpenter, Cilium, and ArgoCD/GitOps, evolve cross-account infrastructure and access, and maintain Terraform/Terragrunt IaC with safe CI-based workflows. Improve deploy, backup/restore, schema-change, and incident-response tooling while reducing operational load and cloud spend.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Turn a fast-growing, stateful system into a predictable, well-automated platform by handling provisioning, scaling, rebalancing, and recovery.
  • •Operate EKS clusters across environments using Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments.
  • •Manage and evolve a multi AWS account organization, including provisioning, networking, access control, and cross-account connectivity.
  • •Maintain the Terraform/Terragrunt IaC platform with modules and automated plan-on-PR/apply-on-merge pipelines using safe shared-infrastructure patterns.
  • •Improve operational tooling for deploys, schema changes, backups/restores, and incident response; participate in on-call/incident response to make incidents rarer over time.

Key Requirements

  • •Deep hands-on experience with Kubernetes in production (EKS preferred), including debugging node pressure, networking issues, and deployment failures at scale.
  • •Strong experience operating production infrastructure on AWS, including understanding organizational boundaries, IAM, and networking across many accounts.
  • •Experience automating infrastructure with Terraform or Terragrunt at scale, including module design and state management.
  • •Solid understanding of Linux systems including disk, memory, networking, and failure modes.
  • •Experience supporting stateful systems (databases, queues, storage systems) and debugging performance/reliability issues in production.
Skills:Deep ownershipProblem-solvingDebuggingCollaborationAutomation mindsetReliability focus
Tech Stack:AWSKubernetesEKSKarpenterCiliumArgoCDGitOpsTerraformTerragruntIaCLinuxVMsGitHub ActionsCI/CD

Company Brief

PostHog
PostHog provides an open-source product analytics and developer platform that combines analytics, session replay, feature flags, A/B testing, and data warehousing to help engineering and product teams build and optimize products.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2020
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor