CoDesign & NextGen Performance Engineer

Cerebras
Sunnyvale, Toronto
Workplace: OnsiteFull timeFunction: Design (Product/UX/UI/Visual)Experience: 3+ yearsEducation: high_schoolSkills: ["Analytical","Problem-solving","Debugging","Performance optimization"]

Characterize, analyze, and optimize the performance of state-of-the-art AI models running on Cerebras’ Wafer Scale Engine hardware. Work across the hardware/software stack to find bottlenecks, improve computational efficiency, and influence Cerebras’ next-generation AI architecture and software systems. Bring up new WSE generations, build kernel-level and end-to-end performance models, optimize kernel micro code and compiler algorithms, and develop tools to visualize performance data across the system and cluster.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 months ago

CoDesign & NextGen Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Characterize, analyze, and optimize the performance of state-of-the-art AI models running on Cerebras’ Wafer Scale Engine hardware. Work across the hardware/software stack to find bottlenecks, improve computational efficiency, and influence Cerebras’ next-generation AI architecture and software systems. Bring up new WSE generations, build kernel-level and end-to-end performance models, optimize kernel micro code and compiler algorithms, and develop tools to visualize performance data across the system and cluster.
Location: Sunnyvale, Toronto
Workplace: Onsite
Employment Type: Full time
Job Function: Design (Product/UX/UI/Visual)
Seniority: Mid level

Key Responsibilities

  • •Bring up and optimize performance on new generations of the Cerebras WSE.
  • •Build kernel-level and end-to-end performance models to estimate performance for state-of-the-art and customer ML models.
  • •Optimize and debug kernel micro code and compiler algorithms to improve inference speed, throughput, and compute utilization.
  • •Debug and understand runtime performance across the system and compute cluster.
  • •Develop tools and infrastructure to visualize performance data from the Wafer Scale Engine and compute cluster.

Key Requirements

  • •Bachelors, Masters, or PhD in Electrical Engineering or Computer Science, with strong background in computer architecture.
  • •Exposure to low-level deep learning/LLM math.
  • •3+ years of experience in areas like computer architecture, CPU/GPU performance, kernel optimization, or HPC.
  • •Experience building performance models, working with kernel-level and end-to-end performance estimation.
  • •Comfort with C++ and Python, including experience with CPU/GPU simulators and performance profiling/debug.
Experience:3+ years
Education:High School in Electrical Engineering or Computer Science
Skills:AnalyticalProblem-solvingDebuggingPerformance optimization
Tech Stack:C++PythonCPU/GPULLMHPC

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn