Distributed Systems Engineer, Data & Inference Platform

Adaption Labs
San Francisco
Workplace: RemoteFull timeFunction: IT Operations (Systems/Network Admin)Experience: 5+ yearsSkills: ["Communication","Problem-solving","Collaboration"]

Build and operate distributed inference systems for large language models and the data pipelines that feed them. You’ll optimize throughput, latency, and cost across GPU fleets, design production-grade Ray Data or Spark pipelines, and own on-call reliability. Collaborate with researchers and ML engineers to take workloads from experimentation to production, ensuring scalable, efficient, and cost-effective inference at petabyte scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Adaption Labs
Adaption Labs
2 months ago

Distributed Systems Engineer, Data & Inference Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build and operate distributed inference systems for large language models and the data pipelines that feed them. You’ll optimize throughput, latency, and cost across GPU fleets, design production-grade Ray Data or Spark pipelines, and own on-call reliability. Collaborate with researchers and ML engineers to take workloads from experimentation to production, ensuring scalable, efficient, and cost-effective inference at petabyte scale.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Sr. Manager level

Key Responsibilities

  • •Serve Models at Scale: Design and operate distributed inference systems for LLMs, optimizing throughput, latency, and cost across heterogeneous GPU fleets. Batching, scheduling, KV cache management, autoscaling — you own the levers that make inference economical.
  • •Move the Data: Build large-scale data pipelines (Ray Data, Spark, or equivalents) that ingest, transform, and curate the datasets behind training and evaluation. The bottleneck is rarely where people think it is, and you find it.
  • •Debug the Undebuggable: Chase down the failure modes that only emerge under real production traffic — stragglers, head-of-line blocking, silent data corruption, GPU memory fragmentation — and write the postmortems that prevent the next ten. Define SLOs, build the observability to measure them, and own the on-call rotation that defends them.
  • •Partner Across the Stack: Work directly with researchers and ML engineers to take experimental workloads from "runs on one node" to "runs in production." You're a systems partner, not a ticket queue.

Pay and Benefits

Perks:Flexible WorkMedical BenefitsMeal AllowancePaid LeaveTravel Allowance

Key Requirements

  • •5+ years building and operating distributed systems in production.
  • •Deep experience with at least one large-scale data or compute framework (Ray, Spark, Flink, Beam, Dask).
  • •Strong fluency in Python and at least one systems language (Go, Rust, C++).
  • •Working knowledge of the GPU/accelerator stack: CUDA fundamentals, NCCL, mixed precision, memory layout.
  • •Experience operating Kubernetes-based infrastructure, including custom operators or schedulers.
Experience:5+ yearsDistributed systemsData pipelinesGPU computingRaySparkKubernetes
Skills:CommunicationProblem-solvingCollaboration
Languages:English
Tech Stack:PythonGoRustC++CUDANCCLRaySparkFlinkBeamDaskKubernetes

Company Brief

Adaption Labs
Adaption Labs delivers AI consulting, engineering, and MLOps services to help organizations design, build, and deploy machine learning and generative AI solutions. They focus on productizing models, integrating AI into production systems, and accelerating digital transformation for enterprise clients.
Industry: Consulting
Website