Member of Technical Staff (AI Infrastructure Engineer)

Perplexity
London
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Communication","Problem-solving","Teamwork","Collaboration"]

Join our AI infra team to design, deploy, and optimize large-scale AI training and inference clusters. You’ll run Kubernetes and Slurm-based HPC environments on AWS, build APIs for training pipelines and inference services, and partner with Inference and Research teams to ensure high uptime, scalable resource utilization, and robust observability for ML workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Perplexity
Perplexity
5 months ago

Member of Technical Staff (AI Infrastructure Engineer)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Join our AI infra team to design, deploy, and optimize large-scale AI training and inference clusters. You’ll run Kubernetes and Slurm-based HPC environments on AWS, build APIs for training pipelines and inference services, and partner with Inference and Research teams to ensure high uptime, scalable resource utilization, and robust observability for ML workloads.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads
  • •Manage and optimize Slurm-based HPC environments for distributed training of large language models
  • •Develop robust APIs and orchestration systems for both training pipelines and inference services
  • •Implement resource scheduling and job management systems across heterogeneous compute environments
  • •Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

Key Requirements

  • •Kubernetes administration, including custom resource definitions, operators, and cluster management
  • •Slurm workload management, including job scheduling, resource allocation, and cluster optimization
  • •Experience with deploying and managing distributed training systems at scale
  • •Deep understanding of container orchestration and distributed systems architecture
  • •High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)
Experience:AIMLHPCAWS
Skills:CommunicationProblem-solvingTeamworkCollaboration
Languages:English
Tech Stack:KubernetesSlurmPythonC++PyTorchAWSGPU

Company Brief

Perplexity
Perplexity AI provides an AI-powered answer engine that returns conversational, citation-backed answers to user queries and offers Pro/Enterprise products and APIs for research, knowledge work, and search augmentation.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2022
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor