Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity
San Francisco, Seattle, New York
Workplace: HybridFull timeUSD 250,000 - 485,000 annuallyFunction: Software EngineeringSkills: ["Ownership","End-to-end problem-solving","Technical partnership","Reliability focus","Distributed systems thinking"]

Own the platform that powers Perplexity’s real-time training and inference workloads on a multi-cloud GPU fleet. Build a self-serve compute platform, operate GPU provisioning and lifecycle management, and design scheduling/placement to address GPU scarcity. Develop Kubernetes-based GPU orchestration (operators, CRDs, multi-cluster federation) with fault tolerance, autoscaling, and observability so both long-running training and low-latency inference run reliably. Partner with inference and cloud engineers to set the technical roadmap.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Perplexity
Perplexity
6 days ago

Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own the platform that powers Perplexity’s real-time training and inference workloads on a multi-cloud GPU fleet. Build a self-serve compute platform, operate GPU provisioning and lifecycle management, and design scheduling/placement to address GPU scarcity. Develop Kubernetes-based GPU orchestration (operators, CRDs, multi-cluster federation) with fault tolerance, autoscaling, and observability so both long-running training and low-latency inference run reliably. Partner with inference and cloud engineers to set the technical roadmap.
Location: San Francisco, Seattle, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build and own a self-serve compute platform for launching training jobs and operating inference services without managing GPU provisioning or provider infrastructure.
  • •Operate the GPU fleet, owning provisioning, lifecycle management, reliability, and capacity integration across providers.
  • •Design scheduling and placement logic to find available capacity, pack it efficiently, and place workloads under constraints.
  • •Support both long-running distributed training and high-availability, low-latency production inference on the same fleet.
  • •Own Kubernetes-based GPU orchestration (operators and CRDs), manage many clusters across providers, and build fault tolerance, autoscaling, and observability.

Pay and Benefits

Salary: USD 250,000 - 485,000 annually

Key Requirements

  • •Deep Kubernetes experience with custom operators, CRDs, and multi-cluster federation.
  • •Experience managing GPU clusters at scale, including NVIDIA hardware, CUDA, and fast networking (InfiniBand or RoCE).
  • •Ability to orchestrate compute across multiple clouds (e.g., CoreWeave, AWS, GCP) and handle provider differences.
  • •Strong distributed systems fundamentals including scheduling, resource allocation, and fault tolerance under load.
  • •Infrastructure/system-level coding experience in Go, Rust, or C++.
Experience:AI inferenceGPU clustersDistributed systemsKubernetesMulti-cloudHPC
Skills:OwnershipEnd-to-end problem-solvingTechnical partnershipReliability focusDistributed systems thinking
Tech Stack:KubernetesCustom operatorsCRDsMulti-cluster federationGoRustC++NVIDIACUDAInfiniBandRoCECoreWeaveAWSGCPVLLMSGLangTensorRT-LLMSlurmCUDA kernelsTriton

Company Brief

Perplexity
Perplexity AI provides an AI-powered answer engine that returns conversational, citation-backed answers to user queries and offers Pro/Enterprise products and APIs for research, knowledge work, and search augmentation.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2022
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor