Infrastructure engineer (UK)

Writer
London
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Communication","Collaboration","Problem-solving","Ownership","Decision-making"]

Own the reliability, performance, and efficiency of mission-critical services for an enterprise AI platform, defining SLOs/error budgets and carrying on-call. Build resilient, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure using Kubernetes, Helm, and Terraform (or Pulumi). Lead incident response with deep root-cause analysis, automate operational toil with Python/Go, and integrate AI-assisted workflows (agents) into day-to-day engineering.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Writer
Writer
16 hours ago

Infrastructure engineer (UK)

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Own the reliability, performance, and efficiency of mission-critical services for an enterprise AI platform, defining SLOs/error budgets and carrying on-call. Build resilient, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure using Kubernetes, Helm, and Terraform (or Pulumi). Lead incident response with deep root-cause analysis, automate operational toil with Python/Go, and integrate AI-assisted workflows (agents) into day-to-day engineering.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own end-to-end reliability, performance, and efficiency of core services, including SLOs/error budgets and on-call coverage.
  • •Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure using Kubernetes, Helm, and Terraform/Pulumi.
  • •Lead incident response, post-mortems, and root-cause analyses, applying learnings to prevent repeat incidents.
  • •Automate operational tasks and infrastructure management with Python or Go, reducing toil and improving release pipelines and observability.
  • •Use AI/agent tooling in daily operations to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review PRs.

Pay and Benefits

Perks:Health InsuranceDentalPensionPaid ParentalLearning BudgetGym Membership

Key Requirements

  • •5+ years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, high-availability production systems at a high-growth product company.
  • •Experience running containerisation in production with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred).
  • •Strong proficiency in Python or Go for automation and infrastructure tooling.
  • •Daily workflow includes agentic AI tooling (e.g., Claude Code, Droid, Codex, internal skills); candidates without AI tooling in their workflow won’t be advanced.
  • •Demonstrated first-principles decision-making to challenge the status quo and propose reliability solutions based on failure modes and tradeoffs.
Experience:5+ yearsEnterpriseSaaSAIHigh-growth
Skills:CommunicationCollaborationProblem-solvingOwnershipDecision-making
Tech Stack:PythonGoAWSGCPAzureKubernetesHelmTerraformPulumiClaude CodeDroidCodexPrometheusGrafanaELKPython or GoRunbooksTerraform/Helm

Company Brief

Writer
Provides an AI writing platform for enterprises that helps teams create on-brand, high-quality content at scale using customizable style guides, real-time suggestions, and governance controls.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series B
Headquarters: New York, United States
Founded: 2019
WebsiteLinkedIn