Senior Platform Engineer AI Infrastructure

MercadoLibre
Bogotá, Buenos Aires
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Troubleshooting","Root cause analysis","Continuous improvement","Strategic decision-making","Engineering best practices"]

Design and scale secure AI/ML compute infrastructure by operating Kubernetes in production and supporting high-throughput workloads. Diagnose performance interactions between infrastructure and AI/ML workloads, investigate distributed-system issues to reach root causes, and continuously improve deployment, scheduling, networking, and observability. Build expertise in cutting-edge distributed computing and GPU performance to lead an advanced regional capability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
MercadoLibre
MercadoLibre
5 days ago

Senior Platform Engineer AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design and scale secure AI/ML compute infrastructure by operating Kubernetes in production and supporting high-throughput workloads. Diagnose performance interactions between infrastructure and AI/ML workloads, investigate distributed-system issues to reach root causes, and continuously improve deployment, scheduling, networking, and observability. Build expertise in cutting-edge distributed computing and GPU performance to lead an advanced regional capability.
Location: Bogotá, Buenos Aires
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate critical compute infrastructure at scale by managing Kubernetes environments in production to support high-throughput AI/ML workloads.
  • •Analyze and diagnose how infrastructure interacts with AI/ML workloads to optimize performance and make strategic technical decisions.
  • •Investigate and resolve complex issues in distributed environments efficiently to identify root causes.
  • •Continuously evolve the platform by proposing and driving improvements in deployment, scheduling, networking, and observability.
  • •Develop and expand expertise in frontier technologies for distributed computing and GPU performance to lead in a regional innovation area.

Key Requirements

  • •Experience operating, managing, troubleshooting, and performing advanced configuration of Kubernetes in production.
  • •Experience administering infrastructure on cloud providers, especially AWS.
  • •Proficiency with infrastructure-as-code tools such as Terraform (or equivalents).
  • •Knowledge and judgment working with observability using metrics, logs, and alerting.
  • •Programming or scripting knowledge for automation and operations using Python, Go, or Bash.
Experience:AI/MLCloud infrastructureDistributed systemsGPUKubernetes
Skills:TroubleshootingRoot cause analysisContinuous improvementStrategic decision-makingEngineering best practices
Tech Stack:KubernetesAWSTerraformObservabilityMetricsLogsAlertingPythonGoBashDeploymentSchedulingNetworkingGPUAI/ML

Company Brief

MercadoLibre
Operates Latin America's leading e‑commerce marketplace and fintech ecosystem, offering online buying and selling, classifieds, payment processing, credit, and logistics services across multiple countries in the region.
Industry: Online Marketplaces
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Buenos Aires, Argentina
Founded: 1999
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn