Staff Software Engineer, Inference Platform

Cerebras
Sunnyvale, Toronto
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 8+ yearsSkills: ["Technical credibility","Communication","Engineering judgment"]

Lead and contribute to Cerebras’ Inference Platform team, owning the orchestration layer that runs inference across datacenter clusters. You’ll design and maintain production software spanning observability, security, networking, and productionization, while shaping platform direction for Kubernetes-based systems. Drive reliability and performance for active-active, low-latency inference, handle critical production issues, and partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable system designs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 months ago

Staff Software Engineer, Inference Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead and contribute to Cerebras’ Inference Platform team, owning the orchestration layer that runs inference across datacenter clusters. You’ll design and maintain production software spanning observability, security, networking, and productionization, while shaping platform direction for Kubernetes-based systems. Drive reliability and performance for active-active, low-latency inference, handle critical production issues, and partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable system designs.
Location: Sunnyvale, Toronto
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, develop, test, and maintain production software covering testing, continuous development, observability, security, networking, debugging, and productionization.
  • •Shape Inference Platform direction, including Kubernetes custom resource definitions, failure domains, service boundaries, and long-term system evolution.
  • •Architect reliability and performance for active-active systems with rapid failover, graceful degradation, clear SLOs, and improvements in latency, throughput, capacity efficiency, and resilience.
  • •Lead on the hardest production issues and cross-system bottlenecks, driving observability, incident response, capacity planning, and post-incident improvements.
  • •Partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable system designs and align on technical decisions.

Key Requirements

  • •8+ years of software engineering experience building and operating large-scale distributed systems or cloud infrastructure.
  • •Deep expertise in distributed systems architecture, ideally with Kubernetes.
  • •Experience making sound architectural decisions for highly available, latency-sensitive systems at scale.
  • •Experience with security, including certificates, TLS, and mTLS.
  • •Strong proficiency in backend or systems languages such as Go or C++, contributing production code directly.
Experience:8+ years
Skills:Technical credibilityCommunicationEngineering judgment
Tech Stack:KubernetesCI/CDTLSMTLSMetricsLoggingTracingAlerting

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn