Distributed Software Engineer

Cerebras
Toronto
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Debugging","Self-directed learning","Cross-team collaboration","Observability mindset","Rigor to verify"]

Own the software that turns Cerebras’ wafer-scale hardware into a reliable, debuggable “cloud.” Build declarative automation for bare-metal networking, OS, and application software; deliver push-button install/upgrade with canary-gated security patching; and develop Kubernetes operators and control-plane services. Monitor multi-tenant workloads with Prometheus/Grafana pipelines, implement failure detection and HA recovery, and expose CLIs/APIs/MCP gateway for users, operators, and AI agents.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
5 days ago

Distributed Software Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own the software that turns Cerebras’ wafer-scale hardware into a reliable, debuggable “cloud.” Build declarative automation for bare-metal networking, OS, and application software; deliver push-button install/upgrade with canary-gated security patching; and develop Kubernetes operators and control-plane services. Monitor multi-tenant workloads with Prometheus/Grafana pipelines, implement failure detection and HA recovery, and expose CLIs/APIs/MCP gateway for users, operators, and AI agents.
Location: Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build declarative, CRD-driven automation for bare-metal networking, OS, and application software across large clusters.
  • •Deliver push-button cluster install, upgrade, and security patching within downtime budgets using canaries.
  • •Develop Kubernetes operators to schedule large inference workloads with resource locks, priority queues, network topology, and health-aware placement.
  • •Implement gRPC control-plane services including authorization, admission webhooks, and quota policy for a multi-tenant fleet.
  • •Create metrics and log pipelines (exporters for wafer-scale systems, servers, and network fabric) on Prometheus and Grafana, and build failure detection, HA control planes, and recovery with CLIs/APIs/MCP gateway exposure.

Key Requirements

  • •5+ years building and operating production distributed systems or infrastructure software.
  • •Production-quality Go and Python.
  • •Real Kubernetes depth: written/debugged controllers and operators; understand CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC.
  • •Strong debugging skills across distributed systems, Linux, and networking.
  • •Prometheus and Grafana practitioner skills, including PromQL, exporter design, alerting, and cardinality discipline.
Experience:5+ yearsDistributed systemsInfrastructure software
Skills:DebuggingSelf-directed learningCross-team collaborationObservability mindsetRigor to verify
Tech Stack:GoPythonKubernetesCRDsOperatorsControllersReconciliation semanticsInformer cachesAdmission webhooksRBACGRPCPrometheusGrafanaPromQLLinuxRedfishIPMIGNMISFlowMCP gateway

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn