Full Stack LLM Engineer

Cerebras
Toronto
Workplace: HybridFull timeFunction: Software EngineeringSkills: []

Join the Inference Core Model Bringup team to rapidly deploy state-of-the-art open-source or customer-proprietary LLMs on Cerebras CSX systems. You’ll work end to end across model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning. The role focuses on debugging performance and correctness across model code, compiler IRs, runtime behavior, and hardware utilization, plus prototyping tool/API/automation improvements to accelerate bring-up.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 year ago

Full Stack LLM Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Join the Inference Core Model Bringup team to rapidly deploy state-of-the-art open-source or customer-proprietary LLMs on Cerebras CSX systems. You’ll work end to end across model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning. The role focuses on debugging performance and correctness across model code, compiler IRs, runtime behavior, and hardware utilization, plus prototyping tool/API/automation improvements to accelerate bring-up.
Location: Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Contribute to the end-to-end bring up of ML models on Cerebras CSX systems.
  • •Work across the stack, including model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
  • •Debug performance and correctness issues across model code, compiler IRs, runtime behavior, and hardware utilization.
  • •Propose and prototype improvements across tools, APIs, or automation flows to accelerate future bring-ups.

Key Requirements

  • •Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field.
  • •Comfort navigating the full AI toolchain, including Python modeling code, compiler IRs, and performance profiling.
  • •Strong debugging skills across performance, numerical accuracy, and runtime integration.
  • •Experience with deep learning frameworks such as PyTorch and TensorFlow, plus familiarity with model internals like attention, MoE, and diffusion.
  • •Proficiency in C/C++ with low-level optimization and proven compiler development experience with LLVM and/or MLIR.
Education:
Tech Stack:PythonLLaMAQwenPyTorchTensorFlowC/C++LLVMMLIRCompiler IRsRuntime integration

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn