Staff GPU Inference SDET

Cerebras
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: QA, Test & Release EngineeringExperience: 8+ yearsSkills: ["Root-cause analysis","Debugging","Stress testing","Performance benchmarking","Fault injection"]

Build the founding quality and validation platform for a new GPU Inference Development team. Design and scale end-to-end release qualification and automated test suites for a distributed GPU inference stack, including multi-node cluster bring-up, LLM serving workloads, continuous batching, caching, and KV-cache efficiency. Validate numerical correctness across precision and hardware changes, run fault-injection for fleet resilience, and integrate CI/CD telemetry to ensure production-grade reliability and peak inference performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 days ago

Staff GPU Inference SDET

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Build the founding quality and validation platform for a new GPU Inference Development team. Design and scale end-to-end release qualification and automated test suites for a distributed GPU inference stack, including multi-node cluster bring-up, LLM serving workloads, continuous batching, caching, and KV-cache efficiency. Validate numerical correctness across precision and hardware changes, run fault-injection for fleet resilience, and integrate CI/CD telemetry to ensure production-grade reliability and peak inference performance.
Location: Sunnyvale, Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: QA, Test & Release Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the GPU inference stack.
  • •Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode performance and workload validation metrics.
  • •Build automated workload replay and benchmarking tools to validate GPU performance models and track TTFT, ITL, throughput, P99 latency, and capacity efficiency.
  • •Create validation infrastructure for numerical correctness, precision stability (FP16/FP8/quantization), determinism, and output correctness across updates.
  • •Engineer chaos engineering and fault-injection suites to simulate failures and verify automated recovery and fleet resilience, with CI/CD and telemetry integration.

Key Requirements

  • •8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer.
  • •Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters across public cloud or enterprise data centers.
  • •Deep understanding of LLM serving engines and distributed runtimes, including prefill vs. decode, KV-cache management, and dynamic batching.
  • •Expert-level Python skills designing custom test automation frameworks, diagnostic tooling, and CI/CD integration.
  • •Proficiency in container orchestration and high-performance interconnects (e.g., Kubernetes, Slurm, Ray; InfiniBand/RoCE/NCCL) and strong failure analysis for distributed systems.
Experience:8+ yearsAILLM servingGPU inferenceDistributed systemsCloud infrastructureInfrastructure testing
Skills:Root-cause analysisDebuggingStress testingPerformance benchmarkingFault injection
Tech Stack:PythonKubernetesSlurmRayInfiniBandRoCENCCLPrometheusGrafanaFP16FP8QuantizationPyTorch ProfilerNVTXROCmHIPC++

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn