Software Engineer, ML Inference Platform

Dialpad
Buenos Aires
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Experience: 6+ yearsSkills: ["Scrappy","Curious","Optimistic","Persistent","Empathetic","Operational judgment","Systems thinking","Collaboration","Performance awareness"]

Build the production inference platform that serves Dialpad’s in-house AI models at scale. You’ll design and improve low-latency, high-throughput inference serving, GPU utilization, and deployment safety using Kubernetes/GCP and NVIDIA GPUs. Collaborate with model developers and engineers to integrate model serving runtimes (vLLM, Triton, TGI), strengthen observability and debugging, and deliver benchmarking, evaluation, traffic management, and rollback-safe releases.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Dialpad
Dialpad
1 month ago

Software Engineer, ML Inference Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build the production inference platform that serves Dialpad’s in-house AI models at scale. You’ll design and improve low-latency, high-throughput inference serving, GPU utilization, and deployment safety using Kubernetes/GCP and NVIDIA GPUs. Collaborate with model developers and engineers to integrate model serving runtimes (vLLM, Triton, TGI), strengthen observability and debugging, and deliver benchmarking, evaluation, traffic management, and rollback-safe releases.
Location: Buenos Aires
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Design, build, and improve systems connecting AI capability development to production inference.
  • •Build model-serving pathways for low-latency, high-throughput, high-availability inference workloads.
  • •Operate and optimize containerized workloads on Kubernetes/GCP for efficient NVIDIA GPU, memory, storage, and networking utilization.
  • •Integrate and adapt model-serving frameworks/runtimes (vLLM, Triton, TGI) to internal deployment, observability, and release requirements.
  • •Implement traffic and release safety (shadow serving, canary rollouts, staged deployments, rollback mechanisms), plus benchmarking/evaluation tooling and observability improvements.

Key Requirements

  • •6+ years of professional software engineering experience shipping backend services, infrastructure systems, or production platforms.
  • •Proficiency writing maintainable production code in Python, Go, or another backend-oriented language with strong debugging and systems thinking.
  • •Experience building, operating, or optimizing high-throughput services, distributed systems, or data/ML infrastructure focused on latency, reliability, and resource utilization.
  • •Hands-on experience with containers, Kubernetes, and Linux, including CI/CD, deployment automation, and production operations.
  • •Strong judgment for reproducibility, observability, rollout safety, failure modes, and overall system resilience.
Experience:6+ yearsAIBackend servicesCloud infrastructureDistributed systemsData/ML infrastructure
Skills:ScrappyCuriousOptimisticPersistentEmpatheticOperational judgmentSystems thinkingCollaborationPerformance awareness
Languages:En
Tech Stack:PythonGoKubernetesLinuxCI/CDContainersGCPNVIDIA GPUsVLLMTritonTGIStructured loggingTracingDashboardsAlertingObservability

Company Brief

Dialpad
Provides cloud-based business communications platform offering voice, video, messaging, contact center, and AI-driven call intelligence to help teams communicate and collaborate across devices and locations.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: San Francisco, United States
Founded: 2011
WebsiteLinkedIn