Site Reliability Engineer

Mistral
New York
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsEducation: mastersSkills: ["Problem-solving","Communication","Collaboration","Self-motivated","Knowledge sharing"]

Shape the reliability, scalability, and performance of production systems and customer-facing applications. Balance day-to-day SRE operations with long-term engineering improvements to reduce operational toil. Design and maintain fault-tolerant infrastructure for web services and ML workloads, build monitoring/incident response, and develop CI/CD and orchestration workflows using modern tools. Collaborate with software and AI/ML research teams to enable safe, reproducible experiments and cloud-agnostic platform capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mistral
Mistral
1 month ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Shape the reliability, scalability, and performance of production systems and customer-facing applications. Balance day-to-day SRE operations with long-term engineering improvements to reduce operational toil. Design and maintain fault-tolerant infrastructure for web services and ML workloads, build monitoring/incident response, and develop CI/CD and orchestration workflows using modern tools. Collaborate with software and AI/ML research teams to enable safe, reproducible experiments and cloud-agnostic platform capabilities.
Location: New York
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain scalable, highly available, fault-tolerant infrastructure for web services and ML workloads.
  • •Ensure platform and training environments are highly available and support replication across HPC clusters.
  • •Operate production systems, troubleshoot incidents, and run occasional on-call rotations with root cause analysis.
  • •Implement and improve monitoring, alerting, and incident response systems to minimize downtime.
  • •Drive infrastructure automation and orchestration improvements using tools like Kubernetes, Flux, and Terraform, and build workflows and tooling for reliability across APIs and training runs.

Pay and Benefits

Perks:Health InsuranceParental LeaveRelocationWellness StipendMeal AllowanceCommuter BenefitsPension

Key Requirements

  • •Master’s degree in Computer Science, Engineering, or a related field.
  • •7+ years of experience in a DevOps/SRE role.
  • •Strong experience with cloud computing and highly available distributed systems.
  • •Experience working on reliability KPIs (observability, alerting, SLAs) and handling site reliability issues in critical environments.
  • •Hands-on experience with CI/CD, containerization, and orchestration tools (Docker, Kubernetes) plus infrastructure-as-code (Terraform or CloudFormation).
Experience:7+ yearsAI/MLDevOps/SREHigh-performance computingHPCML workloadsOpen source
Education:Master's in Computer Science, Engineering or related field
Skills:Problem-solvingCommunicationCollaborationSelf-motivatedKnowledge sharing
Tech Stack:KubernetesFluxTerraformDockerPrometheusGrafanaELK StackDatadogCloudFormationPythonGoBashCI/CDContainerizationOrchestrationMonitoringLoggingAlertingHPC clustersSlurm

Company Brief

Mistral
Develops state-of-the-art large language models and AI systems, offering models and developer tools for natural language understanding, generation, and enterprise AI integrations. Focuses on open research and production-ready model deployments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
WebsiteLinkedIn