Site Reliability Engineer

Mistral
Paris
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsEducation: mastersSkills: ["Problem-solving","Communication","Collaboration"]

Shape the reliability, scalability, and performance of a cloud-agnostic AI platform. Balance production operations with long-term engineering improvements to reduce toil and maximize availability for web services, inference environments, and ML workloads. Build and maintain fault-tolerant infrastructure, implement monitoring and incident response, and develop CI/CD and automation using Kubernetes, Flux, and Terraform while collaborating with software engineers and AI/ML researchers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mistral
Mistral
1 month ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Shape the reliability, scalability, and performance of a cloud-agnostic AI platform. Balance production operations with long-term engineering improvements to reduce toil and maximize availability for web services, inference environments, and ML workloads. Build and maintain fault-tolerant infrastructure, implement monitoring and incident response, and develop CI/CD and automation using Kubernetes, Flux, and Terraform while collaborating with software engineers and AI/ML researchers.
Location: Paris
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain scalable, fault-tolerant infrastructure for web services and ML workloads.
  • •Ensure platform, inference, and training environments are highly available and can be replicated across HPC clusters.
  • •Operate systems in production, troubleshoot issues, and respond to incidents via on-call and root cause analysis.
  • •Implement and improve monitoring, alerting, and incident response to minimize downtime and optimize performance.
  • •Develop and maintain CI/CD, automation, and tooling for containerization, orchestration, monitoring, and logging.

Pay and Benefits

Perks:Health InsuranceParental LeavePensionRelocationWellness StipendMeal AllowanceCommuter Benefits

Key Requirements

  • •Master’s degree in Computer Science, Engineering, or a related field.
  • •7+ years of experience in a DevOps or SRE role with strong expertise in cloud computing and distributed systems.
  • •Hands-on site reliability experience including root cause analysis, in-production troubleshooting, and on-call rotations.
  • •Proficiency with reliability KPIs such as observability, alerting, and SLAs.
  • •Experience with CI/CD, containerization, and orchestration tools such as Docker and Kubernetes.
Experience:7+ yearsAI/MLDevOpsSREDistributed systemsHPC
Education:Master's
Skills:Problem-solvingCommunicationCollaboration
Tech Stack:KubernetesFluxTerraformDockerPrometheusGrafanaELK StackDatadogCloudFormationPythonGoBashCI/CDContainerizationOrchestrationMonitoringLoggingObservability

Company Brief

Mistral
Develops state-of-the-art large language models and AI systems, offering models and developer tools for natural language understanding, generation, and enterprise AI integrations. Focuses on open research and production-ready model deployments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
WebsiteLinkedIn