Software Engineer - Host and Network IO

Cerebras
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: Software EngineeringEducation: mastersSkills: ["Detail-oriented","Clear and effective communication","Meticulous analysis","Rigor","Cross-functional collaboration"]

Build and optimize the Host and Network IO path that connects distributed server nodes to Cerebras’ WSE, including a custom RoCE network stack. You’ll develop control/configuration subsystems, govern a generic IO API, and drive network performance debugging for large AI clusters. Collect and analyze packet traces, develop telemetry tools for visibility, and use kernel bypass and zero-copy techniques to improve CPU/memory utilization. Collaborate across AI IO, cluster architecture, and FPGA/ASIC teams to deliver measurable performance gains.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
14 hours ago

Software Engineer - Host and Network IO

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 43 minutes agoStatus: Live

Job Summary

Build and optimize the Host and Network IO path that connects distributed server nodes to Cerebras’ WSE, including a custom RoCE network stack. You’ll develop control/configuration subsystems, govern a generic IO API, and drive network performance debugging for large AI clusters. Collect and analyze packet traces, develop telemetry tools for visibility, and use kernel bypass and zero-copy techniques to improve CPU/memory utilization. Collaborate across AI IO, cluster architecture, and FPGA/ASIC teams to deliver measurable performance gains.
Location: Sunnyvale, Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop x86 and ARM software to expose next-generation hardware IO capabilities for AI/HPC application teams.
  • •Develop and govern a generic IO API with multiple internal users.
  • •Create control and configuration subsystems that directly interact with Cerebras hardware.
  • •Drive network performance debugging of large AI clusters and root-cause bottlenecks using network statistics and packet traces.
  • •Build tools and telemetry to improve visibility into the network and IO datapath, optimizing CPU/memory utilization and congestion behavior.

Key Requirements

  • •Master’s/PhD in Computer Science or Electrical Engineering with 1+ year industry experience, or 3+ years industry experience.
  • •Experience working in large software environments.
  • •Experience with embedded systems, HW/SW co-design, and some driver development.
  • •Familiarity with network protocols (TCP, RoCE) and network debugging tools such as Wireshark (or willingness to learn).
  • •Some familiarity with network switch environments (e.g., Arista, Juniper) (or willingness to learn).
Education:Master's
Skills:Detail-orientedClear and effective communicationMeticulous analysisRigorCross-functional collaboration
Tech Stack:X86ARMSocket programmingRDMA VerbsRoCETCPWiresharkKernel bypassZero-copyFPGAASICNICPacket traces

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn