Staff Software Engineer, GPU Inference

Cerebras
Toronto, Sunnyvale
Workplace: HybridFull timeFunction: Software EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Technical leadership","Communication","Cross-functional collaboration"]

Build and optimize a GPU inference serving stack for disaggregated AI inference systems, focusing on the full GPU prefill path from API services through vLLM, PyTorch, and the ROCm stack. Drive production reliability with operational readiness, automated recovery, and incident response. Improve latency, throughput, GPU utilization, and memory efficiency, while debugging performance across application, runtime, distributed systems, and hardware layers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 month ago

Staff Software Engineer, GPU Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and optimize a GPU inference serving stack for disaggregated AI inference systems, focusing on the full GPU prefill path from API services through vLLM, PyTorch, and the ROCm stack. Drive production reliability with operational readiness, automated recovery, and incident response. Improve latency, throughput, GPU utilization, and memory efficiency, while debugging performance across application, runtime, distributed systems, and hardware layers.
Location: Toronto, Sunnyvale
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Productionize the GPU inference stack by designing, building, deploying, and maintaining the GPU prefill path across APIs, workers, vLLM, PyTorch, ROCm, GPU nodes, and rack-scale infrastructure.
  • •Own GPU operational readiness including deployment/upgrade/rollback, health-checking, capacity management, failure recovery, and reproducible compatibility for driver/firmware/runtime/model/container.
  • •Drive production reliability by defining service-level indicators, improving fault isolation, graceful degradation, automated recovery, and incident remediation.
  • •Improve inference performance by profiling and optimizing time to first token, throughput, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
  • •Optimize model-serving behavior and debugging across system layers, including scheduling, continuous batching, KV-cache management, quantization, and diagnosing issues across application, runtime, distributed systems, and hardware.

Key Requirements

  • •8+ years of software engineering experience with substantial individual-contributor ownership of complex production systems.
  • •Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or demanding GPU workloads.
  • •Strong programming ability in C++ and Python, including multithreading, concurrency, and performance-sensitive memory management.
  • •Hands-on experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent system.
  • •Experience with Linux and production operations including containers, Kubernetes (or comparable orchestration), observability, CI/CD, and latency-sensitive services.
Experience:8+ yearsAIGPUInference systemsLarge language modelsMultimodalModel serving
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline
Skills:Technical leadershipCommunicationCross-functional collaboration
Tech Stack:C++PythonVLLMSGLangTensorRT-LLMTriton Inference ServerPyTorchROCmHIPLinuxContainersKubernetesCI/CDRDMAKV-cachePrefix cachingContinuous batchingQuantizationGraph executionDistributed communication

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn