Applied Researcher, Audio Post-Training

Cartesia
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 350,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Debugging","Problem-solving","End-to-end ownership","Communication","Experimentation"]

Build and improve capabilities for generative audio models by bridging customer needs and research. Design evaluations to measure new capabilities, create data-processing pipelines to enhance quality, and experiment with finetuning and reinforcement learning to refine model behavior. Root-cause production failures, determine what’s ready for public launch, and translate capability gaps into concrete research plans in collaboration with product and customer stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cartesia
Cartesia
1 month ago

Applied Researcher, Audio Post-Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build and improve capabilities for generative audio models by bridging customer needs and research. Design evaluations to measure new capabilities, create data-processing pipelines to enhance quality, and experiment with finetuning and reinforcement learning to refine model behavior. Root-cause production failures, determine what’s ready for public launch, and translate capability gaps into concrete research plans in collaboration with product and customer stakeholders.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Collaborate with product teams to understand and prioritize customer asks.
  • •Translate vaguely described behavioral problems into concrete research plans.
  • •Ideate and experiment across the modeling stack, including data processing, synthetic data, SFT, RL, and evaluations.
  • •Root-cause failures in production models and apply fixes in future iterations.
  • •Decide which features and capabilities are ready for public launch.

Pay and Benefits

Salary: USD 200,000 - 350,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionParental Leave401kEquityCommuter BenefitsPaid LeaveMeal Allowance

Key Requirements

  • •Strong fundamentals in software engineering and machine learning, including debugging complex systems and a desire to learn quickly.
  • •Experience building and ensuring quality of large multilingual datasets.
  • •Experience training and debugging generative models (speech, text, or multimodal), especially SFT, RL, synthetic data, and human/automated evaluation.
  • •Ability to solve problems grounded in real customer needs rather than just benchmarks.
  • •Bonus: native proficiency in other languages.
Experience:Multilingual datasetsGenerative audioMultimodal
Skills:DebuggingProblem-solvingEnd-to-end ownershipCommunicationExperimentation
Tech Stack:Machine learningSFTReinforcement learningSynthetic dataEvaluation

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Cartesia
Builds real-time multimodal and voice AI (Sonic) that generates expressive, low-latency speech for conversational agents and on-device experiences, serving developers and enterprises with APIs and SDKs.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedInGlassdoor