Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras
Toronto
Workplace: HybridFull timeFunction: QA, Test & Release EngineeringSkills: ["Quick learning","Adaptability","Debugging","Problem solving"]

Build and execute innovative automated test strategies for large-scale AI infrastructure, validating thousands of nodes in high-availability deployments to ensure ~99.9999% reliability. You’ll break distributed ML training/inference and cluster challenges into unit-testable components, develop 100% automation across functional, failure, performance, stress, and security scenarios, and test cluster software (Kubernetes, Prometheus, Grafana) plus key hardware components, championing observability and uptime.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 month ago

Senior Software Development Engineer in Test (SDET) - AI Cluster

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build and execute innovative automated test strategies for large-scale AI infrastructure, validating thousands of nodes in high-availability deployments to ensure ~99.9999% reliability. You’ll break distributed ML training/inference and cluster challenges into unit-testable components, develop 100% automation across functional, failure, performance, stress, and security scenarios, and test cluster software (Kubernetes, Prometheus, Grafana) plus key hardware components, championing observability and uptime.
Location: Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: QA, Test & Release Engineering
Seniority: Mid level

Key Responsibilities

  • •Innovate and execute tests for cutting-edge AI infrastructure by defining optimized test strategies and methodologies.
  • •Develop automated tests (aiming for 100% automation) covering high availability, failure scenarios, performance, stress, and security.
  • •Apply deep understanding of distributed ML training/inference to decompose large system challenges into unit-testable components.
  • •Test AI cluster software components including Kubernetes, Prometheus, and Grafana, plus cluster hardware components such as ML wafer scale accelerators and CPU runtime nodes.
  • •Champion cluster security, reliability/uptime, and ease of use through observability.

Key Requirements

  • •Bachelor's or master's degree in engineering in computer science, electrical, AI, data science or a related field.
  • •5+ years of experience testing enterprise software, distributed systems, and/or datacenter hardware/software.
  • •Strong coding skills in Python, Go (golang), and/or C/C++.
  • •Strong debugging skills for large distributed systems and hardware, including experience with pdb, gdb, strace, and network monitors.
  • •Strong understanding of operating system internals (memory management, file systems, security, performance) and datacenter/device characteristics.
Skills:Quick learningAdaptabilityDebuggingProblem solving
Tech Stack:PythonGoC/C++AWSKubernetesDockerPrometheusGrafanaPdbGdbStrace

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn