Research Engineer, Data Infrastructure

Mistral
Palo Alto, San Francisco
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Experience: 4+ yearsSkills: ["Problem-solving","Debugging","Building scalable systems","Operational excellence","Comfort with ambiguity"]

Build and operate Mistral’s next-generation data infrastructure, scaling massive compute fleets and storage systems for high performance. Help transition from legacy orchestration to decoupled control and data planes, architect multi-cluster orchestration, and design future-proof storage for fine-tuning at extreme scale. Contribute to an internal training platform across Kubernetes and SLURM, implement metadata/lineage systems, and support cloud-native deployments with reliable operations and on-call rotations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mistral
Mistral
1 month ago

Research Engineer, Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate Mistral’s next-generation data infrastructure, scaling massive compute fleets and storage systems for high performance. Help transition from legacy orchestration to decoupled control and data planes, architect multi-cluster orchestration, and design future-proof storage for fine-tuning at extreme scale. Contribute to an internal training platform across Kubernetes and SLURM, implement metadata/lineage systems, and support cloud-native deployments with reliable operations and on-call rotations.
Location: Palo Alto, San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Build and scale massive distributed compute and storage systems for high performance and scalability.
  • •Architect and maintain multi-cluster orchestration layers to optimize workload placement across hardware and regions.
  • •Design future-proof storage formats to support fine-tuning datasets at very large scale.
  • •Develop and improve an internal training platform enabling model training and fine-tuning across Kubernetes and SLURM environments.
  • •Implement metadata and lineage systems and support production-grade pipelines with cloud-native deployment workflows, including on-call rotations.

Key Requirements

  • •4+ years of experience in Data Infrastructure, MLOps, or Infrastructure Engineering.
  • •Experience or strong interest supporting foundational compute and storage platforms.
  • •Proficiency in Python and interest in solving the “brittle data lake” problem with modern, columnar storage standards.
  • •Experience with Kubernetes-native tooling and debugging large-scale distributed systems across multi-cluster environments.
  • •Comfort building and operating scalable, reliable, and secure systems in a rapid-growth AI environment.
Experience:4+ years
Skills:Problem-solvingDebuggingBuilding scalable systemsOperational excellenceComfort with ambiguity
Tech Stack:PythonKubernetesSLURMCloud-nativeDistributed systemsMulti-cluster orchestrationData lakeColumnar storage

Company Brief

Mistral
Develops state-of-the-art large language models and AI systems, offering models and developer tools for natural language understanding, generation, and enterprise AI integrations. Focuses on open research and production-ready model deployments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
WebsiteLinkedIn