AI Inference Engineer - Speech

Zoom Video Communications
Anywhere
Workplace: OnsiteFull timeUSD 151,800 - 332,200 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsEducation: mastersSkills: ["Collaboration","Problem-solving","Performance optimization","Attention to detail","Cross-functional communication"]

Develop and deploy state-of-the-art automatic speech recognition (ASR) services for Zoom products. Optimize ASR inference for production by improving latency, throughput, memory footprint, and resource utilization on NVIDIA GPUs and modern inference hardware (GPU/TPU/AI-specific chips). Collaborate with product, science engineering, and infrastructure teams to build scalable, low-latency, high-accuracy speech model inference solutions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zoom Video Communications
Zoom Video Communications
2 days ago

AI Inference Engineer - Speech

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live
Reposted: similar role first listed 1 day ago

Job Summary

Develop and deploy state-of-the-art automatic speech recognition (ASR) services for Zoom products. Optimize ASR inference for production by improving latency, throughput, memory footprint, and resource utilization on NVIDIA GPUs and modern inference hardware (GPU/TPU/AI-specific chips). Collaborate with product, science engineering, and infrastructure teams to build scalable, low-latency, high-accuracy speech model inference solutions.
Location: Anywhere
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Develop state-of-the-art speech services for Zoom products.
  • •Devise novel techniques when off-the-shelf solutions are not available.
  • •Optimize ASR inference systems for production deployment (latency, throughput, memory footprint, resource utilization).
  • •Improve inference performance via hardware-specific optimization for NVIDIA GPUs.
  • •Profile and debug ASR runtime performance bottlenecks across deployment hardware and environments.

Pay and Benefits

Salary: USD 151,800 - 332,200 annually

Key Requirements

  • •Master's degree in Computer Science, Electrical Engineering, or a related field.
  • •3+ years of experience in speech recognition, speech-LLM, or AI model inference.
  • •Hands-on programming skills in Python, shell scripts, and C/C++ with deep learning fundamentals.
  • •Familiarity with ML frameworks such as PyTorch and TensorFlow and deep understanding of transformer encoder-decoder ASR approaches.
  • •Experience optimizing deep learning inference on NVIDIA GPUs using CUDA, TensorRT, mixed-precision, and custom CUDA kernels (e.g., CUDA Graphs).
Experience:3+ yearsSpeech recognitionSpeech-LLMAI model inferenceAutomatic speech recognitionTransformersReal-time inferenceGPU inference
Education:Master's in Computer Science, Electrical Engineering
Skills:CollaborationProblem-solvingPerformance optimizationAttention to detailCross-functional communication
Tech Stack:PythonShell scriptsC/C++PyTorchTensorFlowTransformer encoder-decoderAttention mechanismsBeam searchSequence-to-sequenceSpeech foundation modelsSpeech-LLMsGPUTPUNVIDIA GPUsCUDATensorRTMixed-precisionCUDA GraphsCUDA kernelsGPU clusters

Company Brief

Zoom Video Communications
Provides video conferencing, chat, phone, webinar, and collaboration software for businesses, schools, and individuals. Its platform is widely used for remote meetings, virtual events, and hybrid work communication.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2011
WebsiteLinkedIn