Cluster Operations Software Engineer

Cerebras
Sunnyvale, Toronto, Bengaluru
Workplace: HybridFull timeFunction: Software EngineeringExperience: 6-8 yearsSkills: ["Problem-solving","Communication","Collaboration","Ownership","Ability to work in a fast-paced environment"]

Manage and operate Cerebras’ AI compute clusters to ensure health, performance, and availability for Wafer-Scale Engine (WSE) workloads. Deploy and troubleshoot container-based services, build operational platforms and reliability tooling, and develop APIs/automation for monitoring, incident response, and fleet management. Work across teams to translate O&M requirements into scalable platform capabilities, with 24/7 monitoring and escalation ownership for distributed systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 day ago

Cluster Operations Software Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Manage and operate Cerebras’ AI compute clusters to ensure health, performance, and availability for Wafer-Scale Engine (WSE) workloads. Deploy and troubleshoot container-based services, build operational platforms and reliability tooling, and develop APIs/automation for monitoring, incident response, and fleet management. Work across teams to translate O&M requirements into scalable platform capabilities, with 24/7 monitoring and escalation ownership for distributed systems.
Location: Sunnyvale, Toronto, Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Deploy, configure, and debug container-based services using Docker.
  • •Build and own software solutions for cluster operations, including monitoring platforms, workflow automation, operational dashboards, and reliability tooling.
  • •Develop APIs, automation services, and integrations to improve operational visibility, incident response, and fleet management across global AI infrastructure.
  • •Monitor and oversee cluster health, proactively identifying and resolving potential issues while maximizing compute capacity through optimization and resource allocation.
  • •Provide 24/7 monitoring and support, perform hands-on troubleshooting, and handle engineering escalations with other teams.

Key Requirements

  • •6-8 years of experience managing and operating complex compute infrastructure, preferably for machine learning or high-performance computing.
  • •Proficiency in Python and Go, including experience building operational platforms, workflow automation, and reliability tooling.
  • •Strong expertise in distributed systems and the ability to troubleshoot and resolve complex technical issues efficiently.
  • •Deep understanding of Linux-based compute systems and command-line tools.
  • •Extensive knowledge of Docker and container orchestration platforms such as k8s.
Experience:6-8 yearsMachine learningHigh-performance computing
Skills:Problem-solvingCommunicationCollaborationOwnershipAbility to work in a fast-paced environment
Tech Stack:PythonGoLinuxDockerK8sKubernetesAPIsAutomationMonitoringAlertingEthernetRoCETCP/IPAWSGCPAzure

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn