Sr. Software Engineer, AI / ML Inference Platform

Dialpad
Buenos Aires
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Mentorship","Technical leadership","Clear communication","Problem-solving","Influence through reasoning"]

Build and improve the shared AI/ML inference platform that turns trained models into reliable, observable, efficient production services. Work end-to-end across GPU training infrastructure, model evaluation and lifecycle tooling, production inference on NVIDIA GPUs in GCP, and operational feedback loops. Partner with ASR and NLP scientists to translate evolving capabilities into scalable systems, balancing latency, reliability, cost, and release safety.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Dialpad
Dialpad
1 day ago

Sr. Software Engineer, AI / ML Inference Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Build and improve the shared AI/ML inference platform that turns trained models into reliable, observable, efficient production services. Work end-to-end across GPU training infrastructure, model evaluation and lifecycle tooling, production inference on NVIDIA GPUs in GCP, and operational feedback loops. Partner with ASR and NLP scientists to translate evolving capabilities into scalable systems, balancing latency, reliability, cost, and release safety.
Location: Buenos Aires
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, build, and improve shared platform capabilities across model training, evaluation, artifact management, release, production inference, and operational feedback.
  • •Build and operate GPU training infrastructure providing reliable, reproducible, and efficient environments for model development and experimentation.
  • •Develop low-latency, high-throughput, highly available inference serving pathways and integrate training/inference frameworks for automation and operational control.
  • •Improve performance and efficiency across compute, memory, storage, networking, batching, concurrency, and workload scheduling for GPU workloads.
  • •Enable safe releases with evaluation checks, shadow traffic, staged rollouts, candidate-vs-incumbent comparisons, fast rollback, and build benchmarking/evaluation and telemetry tooling.

Key Requirements

  • •7+ years of professional software engineering experience with ownership of backend, infrastructure, distributed, or ML platform systems in production.
  • •Experience building or operating systems that support model training, model inference, or the lifecycle connecting them.
  • •Proficiency in Python, Go, or another backend-oriented language, producing maintainable production software with well-designed interfaces.
  • •Hands-on fluency with Linux, containers, Kubernetes, cloud infrastructure, CI/CD, and deployment automation for production operations.
  • •Knowledge of operating GPU workloads and reasoning about utilization, memory, storage, networking, scheduling, and performance.
Experience:ML platformsDistributed systemsGPU workloadsProduction engineeringEnterprise AI
Skills:MentorshipTechnical leadershipClear communicationProblem-solvingInfluence through reasoning
Tech Stack:PythonGoLinuxContainersKubernetesCI/CDGCPGKENVIDIA GPUsVLLMTritonTGIPyTorchJAXObservabilityStructured loggingTracing

Company Brief

Dialpad
Provides cloud-based business communications platform offering voice, video, messaging, contact center, and AI-driven call intelligence to help teams communicate and collaborate across devices and locations.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: San Francisco, United States
Founded: 2011
WebsiteLinkedIn