Founding ML Research Engineer

Kalpa Labs
San Francisco
Workplace: OnsiteFull timeUSD 180,000 - 250,000 annuallyFunction: Research & Scientific (R&D)Experience: 3+ yearsSkills: ["Autonomy","Problem-solving","Debugging","Research curiosity","Ambiguity tolerance"]

Build and ship generalist audio models end-to-end, covering data, pre-training, post-training, and evaluation. Research efficient multi-modal architectures and codecs for speech and audio, then apply SFT/RLHF-style methods, distillation, and preference optimization to enable instruction following and in-context learning across text and audio. Create scalable data and compute pipelines (curation, filtering, mixing, tokenization/feature pipelines) and evaluation harnesses with strong distributed training, reliability, and debugging depth.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Kalpa Labs
Kalpa Labs
2 weeks ago

Founding ML Research Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Build and ship generalist audio models end-to-end, covering data, pre-training, post-training, and evaluation. Research efficient multi-modal architectures and codecs for speech and audio, then apply SFT/RLHF-style methods, distillation, and preference optimization to enable instruction following and in-context learning across text and audio. Create scalable data and compute pipelines (curation, filtering, mixing, tokenization/feature pipelines) and evaluation harnesses with strong distributed training, reliability, and debugging depth.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Research better multi-modal architectures and codecs efficient for spoken speech and general audio.
  • •Post-train audio models for instruction following and in-context learning over both text and audio.
  • •Build large-scale speech model pre-training and post-training pipelines (SFT/RLHF-style, distillation, preference optimization).
  • •Design scalable data + compute pipelines for curation, filtering, mixing, tokenization/feature pipelines, and evaluation harnesses.
  • •Analyze data and conduct extensive audio listening to improve model outcomes.

Pay and Benefits

Salary: USD 180,000 - 250,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Industry/academia experience with pre-training and post-training large neural networks (speech/audio is a plus).
  • •Strong ML systems and engineering depth, including distributed training, performance, and reliability.
  • •Ability to spec, build, debug, and ship while operating in ambiguity.
  • •Experience relevant to speech/audio, and comfort with related modalities like language/vision.
  • •Ability to iterate on research ideas and translate them into production quickly.
Experience:3+ yearsIndustryAcademiaSpeech/audio
Skills:AutonomyProblem-solvingDebuggingResearch curiosityAmbiguity tolerance
Tech Stack:Multi-modalCodecsSFTRLHFDistillationPreference optimizationDistributed trainingData pipelinesEvaluation harnesses

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Kalpa Labs
Develops AI-driven tools and solutions for enterprises, including custom machine learning models, natural language processing, and deployment services to help organizations integrate generative AI and automation into business workflows.
Industry: AI & Machine Learning
Website