AI Infrastructure Operations Engineer

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Data Science & Machine Learning0Education: bachelorsSkills: []

Deploy, bring up, and support Cerebras AI infrastructure in data centers, including CS-X systems, cluster servers, and networking hardware. Execute power-on sequencing, readiness checks, and validation tests, then monitor hardware telemetry and alerts. Triage incidents with first-line troubleshooting, collect logs/telemetry during events, and escalate via established workflows. Build knowledge of Cerebras system architecture and grow toward independent ownership of defined SiteOps workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
7 months ago

AI Infrastructure Operations Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Deploy, bring up, and support Cerebras AI infrastructure in data centers, including CS-X systems, cluster servers, and networking hardware. Execute power-on sequencing, readiness checks, and validation tests, then monitor hardware telemetry and alerts. Triage incidents with first-line troubleshooting, collect logs/telemetry during events, and escalate via established workflows. Build knowledge of Cerebras system architecture and grow toward independent ownership of defined SiteOps workflows.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Entry level

Key Responsibilities

  • •Assist with deployment and bring-up of CS-X systems, cluster servers, and networking hardware, including power-on sequencing, readiness checks, and validation tests.
  • •Monitor hardware telemetry, alerts, and dashboards to support reliable operation and scale-out of AI clusters.
  • •Perform first-line troubleshooting during incidents and execute structured escalation using established workflows.
  • •Collect logs, telemetry, and observations during incidents to support incident response.
  • •Participate in incident response under senior engineer guidance and provide feedback on tooling and process gaps.

Key Requirements

  • •Bachelor’s degree in a relevant engineering field or equivalent experience.
  • •0–3 years experience in hardware operations, systems engineering, or datacenter environments.
  • •Basic familiarity with server hardware and networking fundamentals.
  • •Basic familiarity with Linux systems.
Experience:0Datacenter
Education:Bachelor's
Tech Stack:Linux

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn