Site Reliability Engineer, Mistral Cloud

Mistral
Paris, Amsterdam, Warsaw, Munich, London, Berlin
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: mastersSkills: ["Problem-solving","Communication","Collaboration"]

Shape the reliability, scalability, and performance of the Cloud Platform and customer-facing applications. You’ll design and operate fault-tolerant infrastructure, run production on-call and troubleshooting, and improve monitoring, alerting, and incident response to minimize downtime. Build CI/CD and tooling around containerization and orchestration, drive infrastructure automation, and partner with engineers to enable safe, reproducible model-training experiments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mistral
Mistral
1 month ago

Site Reliability Engineer, Mistral Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Shape the reliability, scalability, and performance of the Cloud Platform and customer-facing applications. You’ll design and operate fault-tolerant infrastructure, run production on-call and troubleshooting, and improve monitoring, alerting, and incident response to minimize downtime. Build CI/CD and tooling around containerization and orchestration, drive infrastructure automation, and partner with engineers to enable safe, reproducible model-training experiments.
Location: Paris, Amsterdam, Warsaw, Munich, London, Berlin
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain scalable, highly available, fault-tolerant infrastructure to support the Cloud platform.
  • •Operate systems and troubleshoot production issues, including on-call responses and infrastructure scaling.
  • •Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • •Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • •Participate in on-call rotations, perform root cause analysis, and drive continuous improvement in infrastructure automation, deployment, and orchestration.

Pay and Benefits

Perks:Health InsuranceParental LeaveRetirementRelocationWellness StipendMeal AllowanceCommuter Benefits

Key Requirements

  • •Master’s degree in Computer Science, Engineering, or a related field.
  • •5+ years of experience in a DevOps or SRE role, with strong expertise in bare metal infrastructure and distributed systems.
  • •Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • •Proficiency working with reliability KPIs such as observability, alerting, and SLAs.
  • •Experience with CI/CD, containerization, and orchestration tools (e.g., Docker and Kubernetes) plus familiarity with Terraform or CloudFormation.
Experience:5+ yearsAI/MLHigh-performance computing (HPC)Distributed systems
Education:Master's
Skills:Problem-solvingCommunicationCollaboration
Tech Stack:PythonGoBashCI/CDDockerKubernetesPrometheusGrafanaELK StackDatadogTerraformCloudFormationContainerizationOrchestrationMonitoringLoggingIncident response

Company Brief

Mistral
Develops state-of-the-art large language models and AI systems, offering models and developer tools for natural language understanding, generation, and enterprise AI integrations. Focuses on open research and production-ready model deployments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
WebsiteLinkedIn