Machine Learning Engineer, Speech LLM Training - San Francisco

Plaud
San Francisco
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["PyTorch","JAX","Distributed training","GPU memory optimization","Signal processing","Foundation models","Inference optimization","Kubernetes","TensorRT-LLM","VLLM","DeepSpeed"]

Lead the development and training of large-scale SpeechLLMs, spanning signal processing to foundation model training. Build novel sequence models, optimize distributed training across GPUs, and push production-ready AI audio capabilities. Deep expertise in PyTorch or JAX, with hands-on experience in large-scale training, memory optimization, and scalable inference in a fast-growing startup environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Plaud
Plaud
2 months ago

Machine Learning Engineer, Speech LLM Training - San Francisco

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Lead the development and training of large-scale SpeechLLMs, spanning signal processing to foundation model training. Build novel sequence models, optimize distributed training across GPUs, and push production-ready AI audio capabilities. Deep expertise in PyTorch or JAX, with hands-on experience in large-scale training, memory optimization, and scalable inference in a fast-growing startup environment.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design and train large-scale SpeechLLMs and related speech models from scratch, spanning ASR, TTS, and generative audio architectures.
  • •Research and implement novel sequence modeling architectures for speech and audio tasks.
  • •Debug and optimize distributed training clusters, ensuring efficient utilization of GPUs and memory.
  • •Develop end-to-end systems from signal processing to deployment, including model training, evaluation, and production-ready inference.
  • •Collaborate across research and engineering teams to push Plaud’s SpeechLLM capabilities into production and scale.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kEquityPaid LeaveHybrid OfficePaid HolidaysParental Leave

Key Requirements

  • •Proven track record building and training large-scale audio or speech models from the ground up (SpeechLLMs, ASR, TTS, or related)
  • •Deep expertise in PyTorch or JAX
  • •Experience optimizing large-scale distributed training, GPU memory utilization, and performance bottlenecks
  • •Comfort traversing the full stack from signal processing to large foundation model training
  • •Ability to work in a fast-paced, high-growth startup and take ownership of ambiguous problems
Experience:SpeechLLMAINLP
Skills:PyTorchJAXDistributed trainingGPU memory optimizationSignal processingFoundation modelsInference optimizationKubernetesTensorRT-LLMVLLMDeepSpeed
Languages:English
Tech Stack:PyTorchJAXDeepSpeedVLLMTensorRT-LLMKubernetesGPUDistributed trainingCUDA

Company Brief

Plaud
Builds AI-native hardware and software note-taking devices and apps (Plaud NOTE, NotePin) that record, transcribe, summarize, and extract insights from conversations to boost professional productivity.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Revenue: USD 100M to 250M
Growth: Scaleup
Funding: Bootstrapped
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn