Member of Technical Staff (AI Inference Engineer)

Perplexity
San Francisco, Palo Alto, New York
Workplace: OnsiteFull timeUSD 210,000 - 385,000 annuallyFunction: Data Science & Machine LearningSkills: ["Problem-solving","Communication"]

Join a growing AI team focusing on large-scale deployment of machine learning models for real-time inference. Work with Python, Rust, C++, PyTorch, Triton, CUDA, and Kubernetes to build robust inference APIs, optimize performance, and improve system reliability across multiple US-based locations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Perplexity
Perplexity
2 years ago

Member of Technical Staff (AI Inference Engineer)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join a growing AI team focusing on large-scale deployment of machine learning models for real-time inference. Work with Python, Rust, C++, PyTorch, Triton, CUDA, and Kubernetes to build robust inference APIs, optimize performance, and improve system reliability across multiple US-based locations.
Location: San Francisco, Palo Alto, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Develop APIs for AI inference that will be used by both internal and external customers
  • •Benchmark and address bottlenecks throughout our inference stack
  • •Improve the reliability and observability of our systems and respond to system outages
  • •Explore novel research and implement LLM inference optimizations

Pay and Benefits

Salary: USD 210,000 - 385,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Experience with ML systems and deep learning frameworks (e.g. PyTorch, TensorFlow, ONNX)
  • •Familiarity with common LLM architectures and inference optimization techniques (e.g. continuous batching, quantization, etc.)
  • •Understanding of GPU architectures or experience with GPU kernel programming using CUDA
Experience:AIMachine learning
Skills:Problem-solvingCommunication
Tech Stack:PythonRustC++PyTorchTritonCUDAKubernetesTensorFlowONNX

Company Brief

Perplexity
Perplexity AI provides an AI-powered answer engine that returns conversational, citation-backed answers to user queries and offers Pro/Enterprise products and APIs for research, knowledge work, and search augmentation.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2022
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor