ML Systems Performance Engineer

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Analytical mindset","Problem-solving mindset"]

Drive end-to-end ML inference performance by building kernel-level and end-to-end performance models, optimizing and debugging kernel microcode and compiler algorithms, and analyzing runtime behavior on the system and cluster. Develop tooling and infrastructure to visualize performance data from the Wafer Scale Engine and compute cluster, enabling performance projection and diagnostics for state-of-the-art and customer models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
7 months ago

ML Systems Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Drive end-to-end ML inference performance by building kernel-level and end-to-end performance models, optimizing and debugging kernel microcode and compiler algorithms, and analyzing runtime behavior on the system and cluster. Develop tooling and infrastructure to visualize performance data from the Wafer Scale Engine and compute cluster, enabling performance projection and diagnostics for state-of-the-art and customer models.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build kernel-level and end-to-end performance models to estimate ML model performance.
  • •Optimize and debug kernel microcode and compiler algorithms to improve inference speed, throughput, and compute utilization.
  • •Debug and understand runtime performance on the system and cluster.
  • •Develop tools and infrastructure to visualize performance data from the Wafer Scale Engine and compute cluster.

Key Requirements

  • •Bachelors, Masters, or PhD in Electrical Engineering or Computer Science.
  • •Strong background in computer architecture.
  • •Exposure to low-level deep learning / LLM math.
  • •3+ years of experience in a relevant domain such as computer architecture, CPU/GPU performance, kernel optimization, or HPC.
  • •Comfort with C++ and Python and experience with CPU/GPU simulators, profiling, and debugging system pipelines.
Experience:3+ yearsHPCCPU/GPU PerformanceComputer ArchitectureKernel Optimization
Education:
Skills:Analytical mindsetProblem-solving mindset
Tech Stack:C++Python

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn