Staff Platform Engineer

DeepL
London
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical ownership","Mentoring","Architecture decision-making","Incident management","Reliability engineering"]

Own the reliability, capacity, and cost efficiency of compute infrastructure powering products and research across AWS and on-prem GPU clusters. Help shape platform architecture and design, build, and run production-grade Kubernetes clusters end-to-end. Drive unification of the hybrid model, define infrastructure-as-code standards, and strengthen observability and security. Mentor engineers, lead incident response, and raise the technical bar across the infrastructure track.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
DeepL
DeepL
1 week ago

Staff Platform Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live
Reposted: similar role first listed 8 months ago

Job Summary

Own the reliability, capacity, and cost efficiency of compute infrastructure powering products and research across AWS and on-prem GPU clusters. Help shape platform architecture and design, build, and run production-grade Kubernetes clusters end-to-end. Drive unification of the hybrid model, define infrastructure-as-code standards, and strengthen observability and security. Mentor engineers, lead incident response, and raise the technical bar across the infrastructure track.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability, capacity, and cost efficiency of compute infrastructure across AWS and on-prem hardware, including hardware lifecycle.
  • •Shape platform architecture and drive technical decisions that serve the best outcome across the track.
  • •Design, build, and operate production-grade Kubernetes clusters across cloud and on-prem, from design to long-term operation.
  • •Unify and deepen the hybrid model so workloads run consistently across on-prem and AWS, supporting research adopting cloud-native ways of working.
  • •Define infrastructure-as-code and platform standards, strengthen observability and security, mentor engineers, and lead incident response for hybrid infrastructure.

Pay and Benefits

Perks:EquityFlexible HoursHybrid WorkAnnual LeaveHealth Insurance

Key Requirements

  • •Hands-on Kubernetes expertise, including designing, building, and operating clusters at scale through the full lifecycle.
  • •Depth in either public cloud or on-prem infrastructure, with credibility to work across both (AWS experience preferred; bare-metal/data-center experience strongly valued).
  • •Networking and Linux depth to debug end-to-end issues from containers to hosts and the network edge.
  • •Infrastructure as code with Terraform (or equivalent) and GitOps delivery using tools such as ArgoCD in production environments you have owned.
  • •Software engineering skills in at least one major language (Go or Python preferred) plus a track record of technical ownership across several teams.
Skills:Technical ownershipMentoringArchitecture decision-makingIncident managementReliability engineering
Tech Stack:KubernetesAWSOn-premGPULinuxTerraformGitOpsArgoCDGoPythonCephBGP

Company Brief

DeepL
DeepL builds Language AI products (DeepL Translator, DeepL Write, APIs and enterprise solutions) that provide high-accuracy translations and writing assistance to businesses and individuals, focusing on privacy, security and enterprise deployment.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Cologne, Germany
Founded: 2017
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor