Member of Technical Staff (AI Inference Engineer)

Perplexity
London
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Self-directed","Problem-solving","Collaborative"]

Join a fast-moving AI team building and running the inference engine behind Perplexity queries. You’ll work on transformer-based models, GPU-accelerated kernels, and a Rust/Python/CUDA/CuTe stack to deploy scalable, low-latency ML inference with strong reliability and observability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Perplexity
Perplexity
9 months ago

Member of Technical Staff (AI Inference Engineer)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join a fast-moving AI team building and running the inference engine behind Perplexity queries. You’ll work on transformer-based models, GPU-accelerated kernels, and a Rust/Python/CUDA/CuTe stack to deploy scalable, low-latency ML inference with strong reliability and observability.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • •GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • •Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • •Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
  • •Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.

Key Requirements

  • •3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • •Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • •Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • •Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
  • •Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
Experience:3+ yearsAIHigh-performance systems
Skills:Self-directedProblem-solvingCollaborative
Languages:English
Tech Stack:RustPythonCUDACuTe DSLTritonCUTLASSPyTorchJAXTensorFlow

Company Brief

Perplexity
Perplexity AI provides an AI-powered answer engine that returns conversational, citation-backed answers to user queries and offers Pro/Enterprise products and APIs for research, knowledge work, and search augmentation.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2022
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor