Inference Engineer

Cartesia
California, San Francisco
Workplace: OnsiteFull timeUSD 180,000 - 250,000 annuallyFunction: QA, Test & Release EngineeringSkills: ["Leadership","Problem-solving","Communication","Collaboration"]

Inference Engineer role focused on designing and building low-latency, scalable model inference and serving stacks for cutting-edge AI foundation models. You’ll collaborate with research and product teams to deliver reliable, cost-effective inference pipelines, build robust infrastructure and monitoring, and shape products across devices with significant autonomy.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cartesia
Cartesia
1 year ago

Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Inference Engineer role focused on designing and building low-latency, scalable model inference and serving stacks for cutting-edge AI foundation models. You’ll collaborate with research and product teams to deliver reliable, cost-effective inference pipelines, build robust infrastructure and monitoring, and shape products across devices with significant autonomy.
Location: California, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: QA, Test & Release Engineering

Key Responsibilities

  • •Design and build low latency, scalable, and reliable model inference and serving stack for our cutting edge foundation models using Transformers, SSMs and hybrid models.
  • •Work closely with our research team and product engineers to serve our suite of products in a fast, cost-effective, and reliable manner.
  • •Design and build robust inference infrastructure and monitoring for our products.

Pay and Benefits

Salary: USD 180,000 - 250,000 annually
Perks:401kRelocation

Key Requirements

  • •Strong engineering skills, comfortable navigating complex codebases and an eye for writing clean and maintainable code.
  • •Experience building large-scale distributed systems with high demands on performance, reliability, and observability.
  • •Technical leadership with the ability to execute and deliver zero-to-one results amidst ambiguity.
  • •Background in or experience working on inference pipelines with machine learning and generative models.
  • •Experience implementing state of the art Machine Learning models and research to applied problems.
Experience:AIMachine LearningMultimodal AI
Skills:LeadershipProblem-solvingCommunicationCollaboration
Tech Stack:TransformersSSMsVLLMSGLangContinuous BatchingCUDATriton

Company Brief

Cartesia
Builds real-time multimodal and voice AI (Sonic) that generates expressive, low-latency speech for conversational agents and on-device experiences, serving developers and enterprises with APIs and SDKs.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedInGlassdoor