Applied Researcher, Audio

Cartesia
California, San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 350,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Problem-solving","Collaboration","Communication"]

Lead research and development of large-scale audio understanding models (multi-speaker ASR, diarization, non-speech audio classification), deploying production-ready systems. Drive self-supervised and few-shot techniques, define evaluation benchmarks, and build pre-training/fine-tuning datasets to advance AI audio perception.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cartesia
Cartesia
11 months ago

Applied Researcher, Audio

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Lead research and development of large-scale audio understanding models (multi-speaker ASR, diarization, non-speech audio classification), deploying production-ready systems. Drive self-supervised and few-shot techniques, define evaluation benchmarks, and build pre-training/fine-tuning datasets to advance AI audio perception.
Location: California, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Architect and develop novel, large-scale models for complex audio understanding tasks, including multi-speaker ASR, diarization, and non-speech audio classification and deploy them to production at scale.
  • •Pioneer research in areas like self-supervised learning for audio, few-shot learning, and robust audio-visual perception.
  • •Set new standards for how we evaluate and benchmark our audio understanding systems.
  • •Build large scale pre-training and fine-tuning datasets for audio understanding capabilities.
  • •Lead high-impact projects that are critical to our mission of building truly aware AI.

Pay and Benefits

Salary: USD 200,000 - 350,000 annually
Equity and Bonus:Equity
Perks:401kHealth InsuranceDentalVision

Key Requirements

  • •Deep expertise in ASR, audio understanding, language modeling, or generative modeling more broadly.
  • •Experience with large-scale training, GPU/TPU acceleration, and model optimization.
  • •Strong applied mindset—able to balance scientific novelty with product impact.
  • •Architect and develop novel, large-scale models for complex audio understanding tasks, including multi-speaker ASR, diarization, and non-speech audio classification and deploy them to production at scale.
  • •Build large scale pre-training and fine-tuning datasets for audio understanding capabilities.
Experience:AIAudio technologyResearch
Skills:Problem-solvingCollaborationCommunication
Tech Stack:ASRDiarizationSelf-supervised learningGPUTPULarge-scale trainingModel optimization

Company Brief

Cartesia
Builds real-time multimodal and voice AI (Sonic) that generates expressive, low-latency speech for conversational agents and on-device experiences, serving developers and enterprises with APIs and SDKs.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedInGlassdoor