Cluster Operations Software Engineer
Sunnyvale, Toronto, Bengaluru
Workplace: HybridFull timeFunction: Software EngineeringExperience: 6-8 yearsSkills: ["Problem-solving","Communication","Collaboration","Ownership","Ability to work in a fast-paced environment"]Manage and operate Cerebras’ AI compute clusters to ensure health, performance, and availability for Wafer-Scale Engine (WSE) workloads. Deploy and troubleshoot container-based services, build operational platforms and reliability tooling, and develop APIs/automation for monitoring, incident response, and fleet management. Work across teams to translate O&M requirements into scalable platform capabilities, with 24/7 monitoring and escalation ownership for distributed systems.
Loading
Loading job details...
Preparing the role view and application actions.

