Agentic Engineer - Voice AI

Wati
Shenzhen
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Problem-solving","Ownership","Bias for action","Proactive mindset","Cross-team collaboration"]

Build and scale real-time voice AI capabilities for WhatsApp, engineering the full pipeline from WebRTC media transport and voice activity detection to LLM inference and text-to-speech. Integrate OpenAI Realtime API and Google Gemini live models, create cascade architectures (ASR → LLM → TTS), and deliver low-latency, scalable infrastructure using LiveKit and RTP/RTCP. Work cross-functionally to ensure reliability and production-grade conversational experiences.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Wati
Wati
3 months ago

Agentic Engineer - Voice AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Build and scale real-time voice AI capabilities for WhatsApp, engineering the full pipeline from WebRTC media transport and voice activity detection to LLM inference and text-to-speech. Integrate OpenAI Realtime API and Google Gemini live models, create cascade architectures (ASR → LLM → TTS), and deliver low-latency, scalable infrastructure using LiveKit and RTP/RTCP. Work cross-functionally to ensure reliability and production-grade conversational experiences.
Location: Shenzhen
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, build, and optimize real-time voice AI pipelines from WebRTC media transport through LLM inference and speech synthesis.
  • •Integrate and orchestrate frontier AI models and cascade architectures (ASR → LLM → TTS) for conversational voice experiences.
  • •Build and maintain media infrastructure including LiveKit-based audio routing, Opus codec handling, RTP/RTCP transport, and voice activity detection.
  • •Develop agent capabilities for voice interactions such as tool calling, function execution, context engineering, and multi-turn conversation management.
  • •Optimize end-to-end latency and ensure reliability, performance, and scalability of voice AI infrastructure for customers.

Key Requirements

  • •3+ years of software engineering experience with strong backend development skills (Go or Python preferred).
  • •Experience with real-time communication technologies such as WebRTC, RTP/RTCP, audio codecs, or media server infrastructure.
  • •Familiarity with AI/LLM integration including model APIs, tool calling, prompt engineering, or agent orchestration.
  • •Experience with speech technologies such as ASR, TTS, voice activity detection, or audio processing pipelines.
  • •Comfort working with PostgreSQL, Redis, and pub/sub messaging systems.
Experience:3+ years
Skills:Problem-solvingOwnershipBias for actionProactive mindsetCross-team collaboration
Tech Stack:GoPythonWebRTCLiveKitOpenAI Realtime APIGoogle Gemini LiveGemini multimodal liveOpusRTP/RTCPVoice activity detectionASRLLMText-to-speechSpeech synthesisTool callingPrompt engineeringPostgreSQLRedisPub/subGCP

Company Brief

Wati
Provides a WhatsApp-based customer engagement platform that helps businesses manage sales, support, and marketing conversations in one shared inbox. The product includes automation, broadcasts, chatbots, and team collaboration tools for messaging-led workflows.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Hong Kong, Hong Kong
Founded: 2018
WebsiteLinkedIn