Senior AI Compute Infrastructure Engineer

Kraken
London
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Python","Linux","Kubernetes","Distributed systems","GPU compute","Infrastructure automation","Cost optimization","Observability","Advanced debugging"]

Join Kraken's AI Compute Infrastructure team to build GPU/accelerator clusters for model training, inference, and experimentation. Design scalable, cost-conscious infrastructure, optimize scheduling and observability, and collaborate with ML researchers to keep AI workloads fast, reliable, and production-grade.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Kraken
Kraken
2 months ago

Senior AI Compute Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Join Kraken's AI Compute Infrastructure team to build GPU/accelerator clusters for model training, inference, and experimentation. Design scalable, cost-conscious infrastructure, optimize scheduling and observability, and collaborate with ML researchers to keep AI workloads fast, reliable, and production-grade.
Location: London
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Own and operate GPU and accelerator clusters used for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration, scheduling primitives, and workload isolation.
  • •Design infrastructure that enables models to run locally on GPUs to reduce external providers and manage compute costs.
  • •Build and improve scheduling, orchestration, placement, quota management, and utilization systems across heterogeneous accelerator environments.
  • •Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using frameworks such as vLLM, Triton Inference Server, TensorRT, or equivalent serving stacks.
  • •Partner with ML engineers and researchers to remove bottlenecks in training, evaluation, batch inference, online inference, deployment, and production debugging workflows.

Key Requirements

  • •5+ years of infrastructure engineering experience with GPU compute, ML infrastructure, distributed systems, or large-scale production platforms.
  • •Hands-on experience operating GPU clusters or accelerator-backed infrastructure in production, including scheduling, orchestration, utilization monitoring, and cost optimization.
  • •Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.
  • •Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.
  • •Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.
Experience:5+ yearsCryptoBlockchainFintech
Skills:PythonLinuxKubernetesDistributed systemsGPU computeInfrastructure automationCost optimizationObservabilityAdvanced debugging
Tech Stack:PythonLinuxKubernetesVLLMTriton Inference ServerTensorRTTorchServeKServeRay ServeCUDA

Company Brief

Kraken
Kraken (Payward, Inc.) is a global cryptocurrency exchange and financial infrastructure provider offering spot and derivatives trading, staking, custody, tokenized assets and institutional services to retail and institutional clients.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2011
Glassdoor
Glassdoor: 4.2
WebsiteLinkedInGlassdoor