Engineering Manager, Kernel Reliability

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Data Analytics & Business IntelligenceExperience: 6+ yearsSkills: ["Technical leadership","Mentoring","Cross-functional collaboration","Execution","Team building"]

Lead a hands-on engineering team focused on kernel-centric reliability for advanced compute clusters and the services that support inference, training, and production workloads. Own the technical vision and roadmap, partner with systems, cluster operations, and debug teams to accelerate failure analysis, and collaborate across SW, ASIC, and hardware architecture to improve reliability and ease of debugging as the system scales.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
8 months ago

Engineering Manager, Kernel Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead a hands-on engineering team focused on kernel-centric reliability for advanced compute clusters and the services that support inference, training, and production workloads. Own the technical vision and roadmap, partner with systems, cluster operations, and debug teams to accelerate failure analysis, and collaborate across SW, ASIC, and hardware architecture to improve reliability and ease of debugging as the system scales.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Manager level

Key Responsibilities

  • •Provide hands-on technical leadership, owning the technical vision and roadmap for kernel-centric reliability across internal and customer-facing systems
  • •Partner with system and cluster operations to reduce downtime after failures using tooling and manual intervention for failure analysis and diagnostics
  • •Collaborate with the debug team to enhance debug tools and speed up failure analysis
  • •Work with SW teams to improve the software stack (including kernels) to strengthen on-field debugging and failure analysis
  • •Lead, mentor, and grow the engineering team while coordinating with ASIC and hardware architecture teams to co-design next-generation architectures for reliability and easier debugging

Key Requirements

  • •6+ years in software engineering, with 3+ years leading teams in SW/HW reliability, debug, diagnostic, failure analysis, or related fields
  • •Expertise in parallel and distributed programming (e.g., message passing, multicore, GPU) and debugging distributed/parallel applications (deadlocks, livelocks, race conditions)
  • •Experience building or using debug and diagnostic tooling (debuggers, core dump handling, code sanitizers)
  • •Deep understanding of computer architectures (instruction pipelining, multithreading, networking)
  • •Background in incident response and post-mortem analysis within monitoring and reliability engineering
Experience:6+ yearsAIDistributed systemsReliability engineeringFailure analysisDebugging
Skills:Technical leadershipMentoringCross-functional collaborationExecutionTeam building
Tech Stack:GPUParallel programmingDistributed programmingMessage passingMulticoreDebuggersCore dump handlingCode sanitizersIncident responsePost-mortem analysis

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn