Member of Technical Staff (AI Infrastructure Engineer)

Perplexity
San Francisco, Palo Alto
Workplace: OnsiteFull timeUSD 210,000 - 385,000 annuallyFunction: Data Science & Machine LearningSkills: ["Teamwork","Problem-solving","Communication"]

Join our AI infrastructure team to design, deploy, and optimize large-scale AI training and inference clusters. You’ll manage Kubernetes and Slurm-based HPC environments on AWS, build APIs for training pipelines, ensure high uptime, and drive autoscaling and observability for ML workloads in production.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Perplexity
Perplexity
1 year ago

Member of Technical Staff (AI Infrastructure Engineer)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join our AI infrastructure team to design, deploy, and optimize large-scale AI training and inference clusters. You’ll manage Kubernetes and Slurm-based HPC environments on AWS, build APIs for training pipelines, ensure high uptime, and drive autoscaling and observability for ML workloads in production.
Location: San Francisco, Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads
  • •Manage and optimize Slurm-based HPC environments for distributed training of large language models
  • •Develop robust APIs and orchestration systems for both training pipelines and inference services
  • •Implement resource scheduling and job management systems across heterogeneous compute environments
  • •Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

Pay and Benefits

Salary: USD 210,000 - 385,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management
  • •Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization
  • •Experience with deploying and managing distributed training systems at scale
  • •Deep understanding of container orchestration and distributed systems architecture
  • •High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)
Experience:AI InfrastructureMLHPCKubernetesAWSSlurm
Skills:TeamworkProblem-solvingCommunication
Tech Stack:KubernetesSlurmPythonC++PyTorchAWSAPIsYAMLObservabilityDocker

Company Brief

Perplexity
Perplexity AI provides an AI-powered answer engine that returns conversational, citation-backed answers to user queries and offers Pro/Enterprise products and APIs for research, knowledge work, and search augmentation.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2022
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor