Machine Learning Engineering Intern

Zendesk
Germany, Portugal
Workplace: RemotePart timeFunction: Data Science & Machine LearningSkills: ["Communication","Listening","Learning mindset"]

Build and improve evaluation systems for Zendesk Explore’s AI Agents by creating golden datasets and frameworks that reflect human judgment. Annotate and curate benchmark conversations, then develop LLM-as-Judge evaluation logic using Python within the Braintrust platform. Compare automated evaluation scores against human feedback and iterate on prompts and logic to ensure metrics directly improve agent quality and behavior.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zendesk
Zendesk
3 days ago

Machine Learning Engineering Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Build and improve evaluation systems for Zendesk Explore’s AI Agents by creating golden datasets and frameworks that reflect human judgment. Annotate and curate benchmark conversations, then develop LLM-as-Judge evaluation logic using Python within the Braintrust platform. Compare automated evaluation scores against human feedback and iterate on prompts and logic to ensure metrics directly improve agent quality and behavior.
Location: Germany, Portugal
Workplace: Remote
Employment Type: Part time
Job Function: Data Science & Machine Learning
Seniority: Intern level

Key Responsibilities

  • •Work with a team building product features for AI Agents.
  • •Create golden datasets by annotating and curating benchmark conversations for training and testing.
  • •Develop LLM-as-Judge evaluation logic in Braintrust using Python to assess AI conversation quality at scale.
  • •Iterate by comparing automated evaluation scores with human feedback.
  • •Refine evaluation prompts and logic until automated metrics improve agent quality and behavior.

Key Requirements

  • •Build datasets and evaluation frameworks by annotating and curating golden datasets from real-world conversations.
  • •Use Python to write and refine LLM-as-Judge evaluation logic within the Braintrust platform.
  • •Have familiarity with Large Language Models (LLMs) and NLP, and show genuine interest in evaluating AI agents.
  • •Be a student and comfortable working across front-end and back-end parts of a SaaS application.
  • •Comfortable in at least one of the main stack languages: Python and TypeScript/JavaScript.
Experience:SaaS
Skills:CommunicationListeningLearning mindset
Tech Stack:PythonTypeScriptJavaScriptLLM-as-JudgeBraintrustNLP

Company Brief

Zendesk
Provides cloud-based customer service and engagement software, including help desk, ticketing, chat, and CRM tools that enable businesses to manage customer support, self-service, and omnichannel communication at scale.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Decacorn (USD 10B+)
Funding: Private Equity Backed
Headquarters: San Francisco, United States
Founded: 2007
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor