Senior Runtime Engineer

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Architectural depth","Low-level implementation","Debugging","Profiling","Collaboration"]

Design and develop high-performance distributed runtime software for training and inference workloads at unprecedented scale. Build and optimize data, communication, and execution pipelines that efficiently utilize CPU, memory, storage, and network across heterogeneous clusters. Collaborate with ML and compiler teams to integrate new model architectures and hardware optimizations, and diagnose performance issues using profiling and instrumentation. Shape system architecture reviews, roadmap planning, and end-to-end model execution efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
10 months ago

Senior Runtime Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design and develop high-performance distributed runtime software for training and inference workloads at unprecedented scale. Build and optimize data, communication, and execution pipelines that efficiently utilize CPU, memory, storage, and network across heterogeneous clusters. Collaborate with ML and compiler teams to integrate new model architectures and hardware optimizations, and diagnose performance issues using profiling and instrumentation. Shape system architecture reviews, roadmap planning, and end-to-end model execution efficiency.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement distributed runtime components for large-scale execution workloads.
  • •Develop and optimize high-performance data and communication pipelines to maximize CPU, memory, storage, and network utilization.
  • •Enable scalable execution across multiple compute nodes with high concurrency and minimal bottlenecks.
  • •Collaborate with ML and compiler teams to integrate new model architectures, training regimes, and hardware-specific optimizations.
  • •Diagnose and resolve complex performance issues across the software stack using profiling and instrumentation tools.

Key Requirements

  • •3+ years developing high-performance or distributed system software.
  • •Strong C/C++ programming skills, including multi-threading, memory management, and performance optimization.
  • •Experience with distributed systems, networking, or inter-process communication.
  • •Solid understanding of data structures, concurrency, and system-level resource management across CPU/I/O/memory.
  • •Bachelor’s, Master’s, or equivalent experience in Computer Science, Electrical Engineering, or a related field.
Experience:3+ yearsAIDistributed systemsMachine learningHPCLarge-model scaling
Education:Bachelor's
Skills:Architectural depthLow-level implementationDebuggingProfilingCollaboration
Tech Stack:C++CPythonPyTorchDistributed systems

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn