Research Engineer, ML Platform

Mistral
Palo Alto
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Experience: 4+ yearsSkills: ["Developer experience focus","Problem diagnosis","Reliability mindset","Ownership","Comfort in ambiguous environments"]

Build and operate the ML platform powering large-scale training, evaluation, and batch inference. You’ll develop APIs and tooling for distributed GPU workloads, design workload orchestration (scheduling, quotas, priorities, preemption, placement), and improve heterogeneous GPU capacity utilization. Partnering across the ML lifecycle, you’ll add observability and reliability mechanisms, create self-service workflows for researchers, and participate in on-call to troubleshoot production systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mistral
Mistral
2 days ago

Research Engineer, ML Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Build and operate the ML platform powering large-scale training, evaluation, and batch inference. You’ll develop APIs and tooling for distributed GPU workloads, design workload orchestration (scheduling, quotas, priorities, preemption, placement), and improve heterogeneous GPU capacity utilization. Partnering across the ML lifecycle, you’ll add observability and reliability mechanisms, create self-service workflows for researchers, and participate in on-call to troubleshoot production systems.
Location: Palo Alto
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.
  • •Orchestrate GPU workloads with queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
  • •Improve provisioning, allocation, and utilization of heterogeneous GPU resources across clusters.
  • •Enable multi-cluster execution based on capacity, data locality, hardware requirements, and organizational priorities.
  • •Operate and harden the platform through observability, failure recovery, capacity planning, and on-call production troubleshooting.

Pay and Benefits

Perks:Health InsuranceParental LeaveRelocationRetirementWellness Stipend

Key Requirements

  • •4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
  • •Proficiency in Python or Go and comfort building production-grade distributed systems.
  • •Strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
  • •Experience with workload scheduling and orchestration concepts such as quotas, priorities, preemption, gang scheduling, and topology-aware placement.
  • •Understanding of distributed ML workloads and GPU infrastructure, including training, fine-tuning, evaluation, checkpointing, batch inference, PyTorch, CUDA, and NCCL.
Experience:4+ yearsDistributed systemsML infrastructureKubernetesGPUFrontier AI
Skills:Developer experience focusProblem diagnosisReliability mindsetOwnershipComfort in ambiguous environments
Tech Stack:PythonGoKubernetesKueueKarpenterVolcanoKyvernoPyTorchCUDANCCLGPUDistributed GPU workloads

Company Brief

Mistral
Develops state-of-the-art large language models and AI systems, offering models and developer tools for natural language understanding, generation, and enterprise AI integrations. Focuses on open research and production-ready model deployments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
WebsiteLinkedIn