Agentic Engineer II - Voice AI

Wati
Shenzhen
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Problem-solving","Strong ownership","Bias for action","Proactive","Debugging"]

Build and scale Wati’s real-time voice AI for WhatsApp, designing low-latency systems that let AI agents listen, think, and speak during live calls. You’ll develop voice AI pipelines from WebRTC media transport through ASR → LLM → TTS, integrate OpenAI Realtime API and Google Gemini live models, and engineer the media infrastructure (LiveKit, RTP/RTCP, codecs, VAD). Contribute to the broader AI agent stack and ensure reliability and scalability globally.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Wati
Wati
2 days ago

Agentic Engineer II - Voice AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Build and scale Wati’s real-time voice AI for WhatsApp, designing low-latency systems that let AI agents listen, think, and speak during live calls. You’ll develop voice AI pipelines from WebRTC media transport through ASR → LLM → TTS, integrate OpenAI Realtime API and Google Gemini live models, and engineer the media infrastructure (LiveKit, RTP/RTCP, codecs, VAD). Contribute to the broader AI agent stack and ensure reliability and scalability globally.
Location: Shenzhen
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, build, and optimize real-time voice AI pipelines from WebRTC media transport to LLM inference and speech synthesis.
  • •Integrate and orchestrate frontier AI models (OpenAI Realtime API, Google Gemini multimodal live) using cascade architectures (ASR → LLM → TTS).
  • •Build and maintain voice media infrastructure including LiveKit audio routing, Opus codec handling, RTP/RTCP transport, and voice activity detection.
  • •Develop agent capabilities for voice interactions such as tool calling, function execution, context engineering, and multi-turn conversation management.
  • •Optimize end-to-end latency across the voice pipeline and collaborate with product/platform teams to deliver production-grade voice AI on WhatsApp.

Key Requirements

  • •3+ years of software engineering experience with strong backend development skills (Go or Python preferred).
  • •Experience with real-time communication technologies such as WebRTC, RTP/RTCP, audio codecs, or media server infrastructure.
  • •Familiarity with AI/LLM integration including model APIs, tool calling, prompt engineering, or agent orchestration.
  • •Experience with speech technologies (ASR, TTS, voice activity detection) and audio processing pipelines.
  • •Understanding of distributed systems, microservices, and cloud-native architectures (GCP preferred), plus comfort with PostgreSQL, Redis, and pub/sub messaging systems.
Experience:3+ yearsAILLMReal-time communication
Skills:Problem-solvingStrong ownershipBias for actionProactiveDebugging
Tech Stack:GoPythonWebRTCLiveKitOpenAI Realtime APIGoogle Gemini LiveGoogle Gemini multimodal liveOpusRTP/RTCPVoice activity detectionASRLLMTTSTool callingFunction executionContext managementPostgreSQLRedisPub/sub messagingGCP

Company Brief

Wati
Provides a WhatsApp-based customer engagement platform that helps businesses manage sales, support, and marketing conversations in one shared inbox. The product includes automation, broadcasts, chatbots, and team collaboration tools for messaging-led workflows.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Hong Kong, Hong Kong
Founded: 2018
WebsiteLinkedIn