Technical Lead Manager, Infrastructure Hardware (Server and Network Systems)

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Program & Project Management (PMO)Education: mastersSkills: ["Multi-team program execution","Risk management","Dependency tracking","Executive-level communication","Operating in ambiguity"]

Drive end-to-end delivery of server and network platform programs across Cerebras CS-3–based AI clusters, from requirements and vendor selection through lab bring-up, qualification, and production rollout. Own integrated schedules, milestones, readiness gates, and risk/change management across OEM/ODM partners, switch/vendors, internal software/runtime teams, QA, and deployment/operations. Lead technical reviews and executive updates while partnering with platform architects to translate architecture into qualification and rollout plans.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
6 months ago

Technical Lead Manager, Infrastructure Hardware (Server and Network Systems)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Drive end-to-end delivery of server and network platform programs across Cerebras CS-3–based AI clusters, from requirements and vendor selection through lab bring-up, qualification, and production rollout. Own integrated schedules, milestones, readiness gates, and risk/change management across OEM/ODM partners, switch/vendors, internal software/runtime teams, QA, and deployment/operations. Lead technical reviews and executive updates while partnering with platform architects to translate architecture into qualification and rollout plans.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Program & Project Management (PMO)
Seniority: Mid level

Key Responsibilities

  • •Own end-to-end program execution for server systems and network equipment in Cerebras clusters, including new platforms, refreshes, and major component/config changes.
  • •Drive requirements gathering into executable plans with clear milestones, readiness gates, and cross-functional deliverables.
  • •Build and manage integrated schedules across vendors and internal teams, tracking dependencies, critical paths, and risks.
  • •Manage OEM/ODM and switch/vendor engagements (RFI/RFP, samples, escalations, roadmap alignment) and coordinate qualification plans, acceptance criteria, and rollout strategies with architects.
  • •Lead qualification and release readiness (lab/staging validation, regression tracking, go/no-go decisions) and own risk/change management into production, including operational readiness alignment with deployment and fleet teams.

Key Requirements

  • •8+ years in Technical Program Management (or similar delivery leadership) for server, network, or infrastructure platforms from concept through production.
  • •B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or equivalent experience.
  • •Experience coordinating complex server and/or datacenter network programs across OEM/ODMs, switch vendors, and internal engineering teams.
  • •Working knowledge of server architecture (CPU/NUMA, memory bandwidth, PCIe, NIC and storage IO) and networking fundamentals (leaf-spine fabrics, switch platforms, high-performance interconnects).
  • •Familiarity with Linux server fleet management, including provisioning, firmware/BIOS, drivers, and field triage.
Experience:AI/MLHPCDistributed infrastructure
Education:Master's
Skills:Multi-team program executionRisk managementDependency trackingExecutive-level communicationOperating in ambiguity
Tech Stack:LinuxFirmwareBIOSDriversLeaf-spine fabricsSwitch platformsCPU/NUMAPCIeNICStorage IORFI/RFPOEM/ODMRegression tracking

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn