Member of Technical Staff - Inference

Prime Intellect
United States
Workplace: RemoteFull timeFunction: OtherExperience: 3+ yearsSkills: ["Communication","Collaboration","Problem-solving"]

Hybrid, remote-friendly role building and optimizing large-scale ML/LLM infrastructure for multi-tenant serving and RL integration. You will develop LLM serving platforms, GPU-aware scheduling, resilience features, autoscaling, and model distribution, while refining inference frameworks, parallelism strategies, and end-to-end performance within Prime Intellect’s open superintelligence stack.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
11 months ago

Member of Technical Staff - Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Hybrid, remote-friendly role building and optimizing large-scale ML/LLM infrastructure for multi-tenant serving and RL integration. You will develop LLM serving platforms, GPU-aware scheduling, resilience features, autoscaling, and model distribution, while refining inference frameworks, parallelism strategies, and end-to-end performance within Prime Intellect’s open superintelligence stack.
Location: United States
Workplace: Remote
Employment Type: Full time

Key Responsibilities

  • •Build a multi-tenant LLM serving platform that operates across our cloud GPU fleets.
  • •Design placement and scheduling algorithms for heterogeneous accelerators.
  • •Implement multi-region/zone failover and traffic shifting for resilience and cost control.
  • •Build autoscaling, routing, and load balancing to meet throughput/latency SLOs.
  • •Optimize model distribution and cold-start times across clusters.

Pay and Benefits

Equity and Bonus:Equity
Perks:Remote WorkEquityVisa SponsorshipOff-sites

Key Requirements

  • •3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.
  • •Hands-on with at least one of vLLM, SGLang, TensorRT-LLM.
  • •Distributed serving infrastructure familiarity (e.g., NVIDIA Dynamo).
  • •Deep understanding of prefill vs decode, KV-cache behavior, batching, sampling, speculative decoding, parallelism strategies.
  • •Full-stack debugging including CUDA/NCCL, drivers/kernels, containers, service mesh/networking, and storage.
Experience:3+ yearsAIMLLLMInferenceCloud
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonPyTorchAWSGCPKubernetesCUDANCCLInfiniBandTensorRTVLLMSGLang

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn