Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

Plaud
San Francisco
Workplace: HybridFull timeUSD 180,000 - 270,000 annuallyFunction: Data Science & Machine LearningSkills: ["Python","CUDA","GPU","GPUs","TensorRT","Triton","VLLM","SGLang","Kubernetes","WebSockets","WebRTC"]

Build and optimize high-throughput, ultra-low-latency inference engines for large language models and speech systems. Collaborate with ML training and backend infra teams to maximize latency/throughput, KV cache performance, and real-time streaming efficiency on multi-GPU clusters, in a hybrid, San Francisco-based setting.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Plaud
Plaud
2 months ago

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Build and optimize high-throughput, ultra-low-latency inference engines for large language models and speech systems. Collaborate with ML training and backend infra teams to maximize latency/throughput, KV cache performance, and real-time streaming efficiency on multi-GPU clusters, in a hybrid, San Francisco-based setting.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design, implement, and optimize inference serving pipelines for real-time AI workloads
  • •Tune latency-throughput tradeoffs and Time-To-First-Token/Time-To-First-Audio in streaming settings
  • •Collaborate with ML training and backend infra teams to align hardware/software across the stack
  • •Scale multi-GPU multi-node inference pipelines and manage autoscaling with Kubernetes
  • •Deliver production-ready, robust, and well-documented systems for global users

Pay and Benefits

Salary: USD 180,000 - 270,000 annually
Equity and Bonus:Equity
Perks:401kHealth InsuranceDentalVisionPaid LeaveEquity

Key Requirements

  • •Hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models
Experience:AISaaS
Skills:PythonCUDAGPUGPUsTensorRTTritonVLLMSGLangKubernetesWebSocketsWebRTC
Tech Stack:PythonCUDANVIDIAGPUsTensorRTNVIDIA TritonVLLMSGLangKubernetesWebSocketsWebRTC

Company Brief

Plaud
Builds AI-native hardware and software note-taking devices and apps (Plaud NOTE, NotePin) that record, transcribe, summarize, and extract insights from conversations to boost professional productivity.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Revenue: USD 100M to 250M
Growth: Scaleup
Funding: Bootstrapped
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn