Software Engineer, Kernel Reliability

Cerebras
United States, Canada
Full timeFunction: Software EngineeringSkills: ["Root-cause analysis","Debugging","Failure analysis","Incident response","Collaboration"]

Join the on-field Kernel Reliability team to improve reliability of compute clusters and the inference, training, and internal production services. You’ll contribute to a kernel-centric reliability roadmap, partner with system and cluster operations to reduce downtime through tooling and hands-on debugging, enhance debug tools for faster failure analysis, and collaborate across software and hardware teams to design next-generation, easier-to-debug architectures. Participate in incident response, RCA, and post-mortems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
12 hours ago

Software Engineer, Kernel Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join the on-field Kernel Reliability team to improve reliability of compute clusters and the inference, training, and internal production services. You’ll contribute to a kernel-centric reliability roadmap, partner with system and cluster operations to reduce downtime through tooling and hands-on debugging, enhance debug tools for faster failure analysis, and collaborate across software and hardware teams to design next-generation, easier-to-debug architectures. Participate in incident response, RCA, and post-mortems.
Location: United States, Canada
Employment Type: Full time
Job Function: Software Engineering
Seniority: Entry level

Key Responsibilities

  • •Contribute to the technical roadmap and execution for kernel-centric reliability of internal and customer-facing systems.
  • •Partner with System and Cluster Operations to reduce downtime after failures through tooling, analysis, and hands-on debugging support.
  • •Enhance debug tools to speed up failure analysis with the Debug Team.
  • •Collaborate with software teams to improve the software stack (including kernels) for better on-field debugging and failure analysis.
  • •Participate in incident response, root-cause analysis, and post-mortems; drive follow-ups that measurably improve reliability over time.

Key Requirements

  • •Strong programming skills in C/C++ and Python.
  • •Solid foundations in operating systems, computer architecture, and systems programming.
  • •Ability to debug complex issues using logs, traces, and standard debugging workflows with interest in root-cause analysis.
  • •Required or demonstrated through projects, internships, or coursework.
  • •Open to new college graduates (as stated).
Skills:Root-cause analysisDebuggingFailure analysisIncident responseCollaboration
Tech Stack:CC++PythonOperating systemsComputer architectureMessage passingParallel programmingMulticoreGPUEmbeddedLogsTraces

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn