Senior Software Engineer - AI Agents

MercadoLibre
Buenos Aires
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Curiosity","Autonomy","Problem-solving","Teamwork"]

Design and scale an evaluation platform that continuously assesses AI agents using production conversations and configurable rubrics. Build and maintain the continuous-evaluation pipeline, support full lifecycle auditing of agents and their assets, and feed results back into model selection and prompt tuning. Create evaluation strategies for complex conversational agents, work with product teams to translate use cases into measurable criteria, and help improve availability and reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
MercadoLibre
MercadoLibre
1 day ago

Senior Software Engineer - AI Agents

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live
Reposted: similar role first listed 1 week ago

Job Summary

Design and scale an evaluation platform that continuously assesses AI agents using production conversations and configurable rubrics. Build and maintain the continuous-evaluation pipeline, support full lifecycle auditing of agents and their assets, and feed results back into model selection and prompt tuning. Create evaluation strategies for complex conversational agents, work with product teams to translate use cases into measurable criteria, and help improve availability and reliability.
Location: Buenos Aires
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build and maintain a continuous evaluation pipeline that collects production conversations and evaluates them programmatically against configurable rubrics.
  • •Evolve the continuous evaluation platform to audit agents and their assets across the full lifecycle.
  • •Integrate the platform into continuous improvement by feeding back into model selection and prompt tuning.
  • •Design evaluation strategies and rubrics for complex conversational agents in high-availability environments.
  • •Collaborate with product teams to translate use cases into measurable evaluation criteria.

Key Requirements

  • •Strong knowledge of languages such as Python or Go and experience designing RESTful APIs.
  • •Familiarity with microservices architectures, distributed systems, and asynchronous processing.
  • •Knowledge of relational databases such as MySQL.
  • •Familiarity with LLMs, their APIs, and evaluation frameworks such as LLM-as-judge.
  • •Curiosity and autonomy to investigate and solve problems, with the ability to work effectively in a team.
Experience:LLMsAI agentsMachine learning
Skills:CuriosityAutonomyProblem-solvingTeamwork
Tech Stack:PythonGoRESTful APIsMicroservicesDistributed systemsAsynchronous processingMySQLLLMsLLM-as-judge

Company Brief

MercadoLibre
Operates Latin America's leading e‑commerce marketplace and fintech ecosystem, offering online buying and selling, classifieds, payment processing, credit, and logistics services across multiple countries in the region.
Industry: Online Marketplaces
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Buenos Aires, Argentina
Founded: 1999
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn