Principal Engineer, Inference Cloud

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: Communications, PR & CommunityExperience: 10+ yearsSkills: ["Technical credibility","Communication","Judgment","Mentorship","Operational rigor"]

Own the Inference Cloud Platform, the cloud layer behind an AI inference service. Drive long-term architecture and direction across multi-region topology, failure domains, and service boundaries. Build and evolve highly reliable, latency-sensitive systems with active-active patterns, graceful degradation, and clear SLOs. Contribute production code on critical paths, lead difficult production issues, and mentor teams through design reviews, standards, and operational rigor.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
11 months ago

Principal Engineer, Inference Cloud

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own the Inference Cloud Platform, the cloud layer behind an AI inference service. Drive long-term architecture and direction across multi-region topology, failure domains, and service boundaries. Build and evolve highly reliable, latency-sensitive systems with active-active patterns, graceful degradation, and clear SLOs. Contribute production code on critical paths, lead difficult production issues, and mentor teams through design reviews, standards, and operational rigor.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Communications, PR & Community
Seniority: Mid level

Key Responsibilities

  • •Define and prioritize the most important platform problems, making tradeoffs on what the platform will and won’t support.
  • •Set the long-term technical direction for the Inference Cloud Platform, including multi-region topology, failure domains, and service boundaries.
  • •Architect reliability and performance improvements using active-active designs, rapid failover, and graceful degradation with clear SLOs.
  • •Contribute production code on critical paths and perform design reviews, making build-vs-buy and other architectural decisions.
  • •Lead hardest production issues and drive observability, incident response, capacity planning, and post-incident improvements across systems.

Key Requirements

  • •10+ years of software engineering experience building and operating large-scale distributed systems or cloud infrastructure.
  • •Deep expertise in distributed systems architecture in cloud environments, including networking, compute orchestration, container platforms, and multi-region production services.
  • •Proven architectural judgment for highly available, latency-sensitive systems at scale.
  • •Experience optimizing latency, throughput, and efficiency in high-QPS systems, including TTFT and tail-latency reduction.
  • •Strong backend/systems proficiency in Go, C++, or Python, with the ability to contribute production code directly.
Experience:10+ yearsDistributed systemsCloud infrastructureAI inferenceModel serving
Skills:Technical credibilityCommunicationJudgmentMentorshipOperational rigor
Tech Stack:GoC++PythonDistributed systemsCloud infrastructureNetworkingCompute orchestrationContainer platformsMulti-regionActive-active systemsCircuit breakingBackpressureLoad sheddingObservabilityMetricsLoggingTracingAlertingSLISLO

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn