Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco

Plaud
San Francisco
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Communication","Collaboration","Problem-solving"]

Join Plaud’s AI R&D team to build scalable evaluation systems and dashboards for Speech LLMs. You’ll collaborate with ML researchers to define metrics, develop Python-based evaluation harnesses, and monitor model health across training runs, enabling reliable benchmarks and rapid debugging in a distributed, security-conscious environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Plaud
Plaud
2 months ago

Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Join Plaud’s AI R&D team to build scalable evaluation systems and dashboards for Speech LLMs. You’ll collaborate with ML researchers to define metrics, develop Python-based evaluation harnesses, and monitor model health across training runs, enabling reliable benchmarks and rapid debugging in a distributed, security-conscious environment.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Build and maintain evaluation harnesses and pipelines to run at scale against live model checkpoints.
  • •Partner with ML researchers to define metrics and benchmarks for Speech LLM performance (e.g., ASR robustness, TTS steering).
  • •Develop dashboards and tooling to monitor model health during training and track key signals.
  • •Debug mid-training anomalies to determine whether issues are architectural, data-related, or infrastructural.
  • •Communicate statistical results and model behaviors clearly to both technical and non-technical stakeholders.

Pay and Benefits

Perks:Health InsuranceDentalVision401kEquityPaid Leave

Key Requirements

  • •Strong software engineering skills (especially in Python) with experience building reliable distributed systems, data pipelines, or evaluation harnesses that can run at scale against live model checkpoints.
  • •Ability to partner with ML researchers to define what 'good' looks like for a Speech LLM and translate capabilities into measurable benchmarks.
  • •Experience building dashboards and tooling to track model health during training, improve signal-to-noise, and reduce evaluation latency.
  • •Ability to rapidly debug anomalous mid-training results to determine if performance drops stem from model architecture, data, or infrastructure.
  • •Excellent communication skills to convey complex statistical results and model behaviors to technical and non-technical stakeholders.
Experience:AIMLSpeechLLMEvaluation
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonDistributed systemsData pipelinesEvaluation harnessesDashboardsWeights & BiasesMLflow

Company Brief

Plaud
Builds AI-native hardware and software note-taking devices and apps (Plaud NOTE, NotePin) that record, transcribe, summarize, and extract insights from conversations to boost professional productivity.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Revenue: USD 100M to 250M
Growth: Scaleup
Funding: Bootstrapped
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn