Software Engineer, Cluster Deployment

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 2+ yearsSkills: ["Problem-solving","Curiosity","Hands-on mindset","Willingness to learn"]

Build and maintain automation that deploys, validates, and manages AI compute clusters across data centers. Work on pushbutton tooling to make large-scale deployments faster, safer, and more reproducible, turning manual steps into tested workflows. Troubleshoot issues spanning Linux, bare-metal, networking, storage, and Kubernetes, and improve reliability with health checks and observability. Contribute to infrastructure-as-code and GitOps workflows with Terraform and Ansible.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 month ago

Software Engineer, Cluster Deployment

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and maintain automation that deploys, validates, and manages AI compute clusters across data centers. Work on pushbutton tooling to make large-scale deployments faster, safer, and more reproducible, turning manual steps into tested workflows. Troubleshoot issues spanning Linux, bare-metal, networking, storage, and Kubernetes, and improve reliability with health checks and observability. Contribute to infrastructure-as-code and GitOps workflows with Terraform and Ansible.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop and maintain automation for deployment workflows, including provisioning, configuration, validation, and operational handoff.
  • •Transform manual deployment steps into tested, repeatable pushbutton workflows.
  • •Participate in hands-on cluster deployments to develop debugging and operational expertise.
  • •Troubleshoot issues across Linux systems, bare-metal servers, networking, storage, Kubernetes, and connectivity.
  • •Contribute to infrastructure-as-code and GitOps workflows (Terraform, Ansible, PR-based change control) and improve reliability with health checks and observability.

Key Requirements

  • •2+ years of mid- to large-scale data center deployment experience.
  • •Strong fundamentals in Python and Bash, with the ability to write scripts and small programs.
  • •Basic Linux experience, including command-line usage, processes, filesystems, and disk troubleshooting.
  • •Working knowledge of Git, including branching, commits, pull requests, and code review.
  • •CS, ECE, or related technical degree (or equivalent practical experience) with curiosity and strong problem-solving skills.
Experience:2+ yearsData centersBare-metalAI infrastructure
Skills:Problem-solvingCuriosityHands-on mindsetWillingness to learn
Tech Stack:PythonBashAnsibleLinuxGitKubernetesTerraformGitOpsPull requestsCode reviewPrometheusGrafanaVLANsBGPArista EOSPXEDHCPIPXERedfishIPMI

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn