Senior DevOps Engineer

Runware
United Kingdom
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureEducation: mastersSkills: ["Communication","Problem solving","Teamwork","Autonomy"]

We are building a GPU-focused, real-time AI inference platform and need a Senior/Staff DevOps Engineer to design, operate, and scale infrastructure across bare-metal, GPUs, and cloud-native environments. You’ll drive automation, CI/CD, monitoring, and security for high-availability, low-latency systems, enabling rapid model launches and reliable performance for millions of users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Runware
Runware
3 months ago

Senior DevOps Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

We are building a GPU-focused, real-time AI inference platform and need a Senior/Staff DevOps Engineer to design, operate, and scale infrastructure across bare-metal, GPUs, and cloud-native environments. You’ll drive automation, CI/CD, monitoring, and security for high-availability, low-latency systems, enabling rapid model launches and reliable performance for millions of users.
Location: United Kingdom
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and scale infrastructure powering real-time AI inference across GPU fleets, bare-metal servers, serverless and containerised production systems
  • •Evolve platform toward elastic, on-demand infrastructure that scales with customer traffic and model demand
  • •Improve critical paths behind request entrypoints, inference services, queues, storage, load balancers and networking
  • •Automate infrastructure operations from provisioning to CI/CD, deployment safety, progressive rollouts and rapid rollback
  • •Build observability for a high-performance AI platform with signals to spot issues early and understand capacity

Pay and Benefits

Equity and Bonus:Equity
Perks:Remote WorkEquity

Key Requirements

  • •Strong experience as a DevOps/SRE/Infrastructure/Platform Engineer with production-scale systems
  • •Deep Linux knowledge and ability to debug real production issues across networking, storage, performance, and services
  • •Hands-on experience with Infrastructure-as-Code, CI/CD pipelines, and deployment workflows
  • •Experience operating high-availability, low-latency, or high-throughput platforms impacting customers
  • •Strong networking fundamentals across TCP/IP, DNS, load balancing, routing, firewalls, proxies, TLS, and HTTP
Experience:Information technologyAI infrastructureGPU
Education:Master's
Skills:CommunicationProblem solvingTeamworkAutonomy
Tech Stack:LinuxInfrastructure-as-CodeCI/CDGPUNVIDIA driversCUDAContainer runtimesVLLMTensorRTTritonLoad balancersNetworkingTLSHTTP

Company Brief

Runware
Runware.ai offers a platform for building, deploying, and managing LLM-powered applications, focusing on orchestration, observability, and operational tooling to ensure reliable, scalable, and safe production AI workflows.
Industry: Developer Tools
Website