Infrastructure engineer

Writer
New York, San Francisco, Seattle, London
Workplace: HybridFull timeUSD 155,000 - 240,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Communication","Collaboration","Problem-solving","Ownership","Decision-making"]

Own end-to-end reliability, performance, and efficiency for mission-critical services supporting enterprise AI workflows. Build and automate resilient, fault-tolerant infrastructure across AWS (preferred) plus GCP and Azure, using Kubernetes, Helm, and Terraform (or Pulumi). Lead incident response and root-cause analysis, define SLOs and error budgets, and drive reliability improvements via observability, release automation, and AI-assisted operational tooling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Writer
Writer
17 hours ago

Infrastructure engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Own end-to-end reliability, performance, and efficiency for mission-critical services supporting enterprise AI workflows. Build and automate resilient, fault-tolerant infrastructure across AWS (preferred) plus GCP and Azure, using Kubernetes, Helm, and Terraform (or Pulumi). Lead incident response and root-cause analysis, define SLOs and error budgets, and drive reliability improvements via observability, release automation, and AI-assisted operational tooling.
Location: New York, San Francisco, Seattle, London
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability, performance, and efficiency end-to-end by defining SLOs and error budgets and carrying on-call responsibility.
  • •Build and automate resilient infrastructure across AWS (preferred), GCP, and Azure, working with Kubernetes, Helm, and Terraform.
  • •Lead incident response, post-mortems, and root-cause analyses, then apply learnings to prevent repeat incidents.
  • •Design and operate release and multi-region infrastructure layouts, including rollback practices to reduce blast radius.
  • •Collaborate with product, security, and engineering to connect reliability work to product and revenue context and guide system design.

Pay and Benefits

Salary: USD 155,000 - 240,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveHealth InsuranceDentalVisionParental LeaveLearning Budget401kEquityWellness Stipend

Key Requirements

  • •5+ years of experience in infrastructure engineering, DevOps, or building/operating large-scale, high-availability production systems.
  • •Hands-on containerisation experience in production, with Helm and Terraform (or Pulumi) on at least one major cloud (AWS preferred).
  • •Strong automation skills in Python or Go for operational tasks and infrastructure tooling.
  • •Use AI-assisted/agentic tooling in your day-to-day workflow (e.g., Claude Code, Droid, Codex) and have strong opinions on reliability.
  • •Demonstrated ability to analyze failure modes, challenge the status quo, and design reliable solutions with clear tradeoffs.
Experience:5+ years
Skills:CommunicationCollaborationProblem-solvingOwnershipDecision-making
Tech Stack:PythonGoAWSGCPAzureKubernetesHelmTerraformPulumiClaude CodeDroidCodexPrometheusGrafanaELK

Company Brief

Writer
Provides an AI writing platform for enterprises that helps teams create on-brand, high-quality content at scale using customizable style guides, real-time suggestions, and governance controls.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series B
Headquarters: New York, United States
Founded: 2019
WebsiteLinkedIn