Audio Inference Engineer, Model Efficiency

Cohere
New York, San Francisco, Toronto, Montreal
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Collaboration","Problem-solving","Communication","Bias for action"]

Join Cohere as an Audio Inference Engineer on the Model Efficiency team to optimize real-time audio inference, reduce latency, and boost throughput. Collaborate with training and serving infra to ensure seamless model deployment, while tackling streaming audio workloads and delivering high-performance solutions for frontier ML models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
10 months ago

Audio Inference Engineer, Model Efficiency

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Join Cohere as an Audio Inference Engineer on the Model Efficiency team to optimize real-time audio inference, reduce latency, and boost throughput. Collaborate with training and serving infra to ensure seamless model deployment, while tackling streaming audio workloads and delivering high-performance solutions for frontier ML models.
Location: New York, San Francisco, Toronto, Montreal
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Advance core audio model serving metrics, including latency, throughput, and quality.
  • •Dive deep into systems to identify bottlenecks and deliver creative solutions for audio processing and streaming workloads.
  • •Collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment.
  • •Focus on real-time and streaming audio inference to improve end-to-end performance.
  • •Contribute to high-performance inference systems and optimize GPU-based workflows.

Pay and Benefits

Perks:Health InsuranceDentalParental LeaveRemote WorkMeal AllowancePaid Leave

Key Requirements

  • •Significant experience developing high-performance audio or machine learning inference systems.
  • •Proficiency with programming languages such as C++ and Python.
  • •Hands-on experience with deep learning models for audio, speech, or language applications.
  • •A bias for action and a strong results-oriented mindset.
  • •Experience with real-time and streaming audio inference is a plus.
Skills:CollaborationProblem-solvingCommunicationBias for action
Tech Stack:C++PythonPyTorchTensorFlowVLLMSGLangTensort-LLMGPU programmingDuplex real-time streamingAudio frameworks

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor