Staff Software Engineer, Inference Cloud

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Architectural judgment","Technical leadership","Debugging","Communication","Mentorship"]

Own major architectural areas of the Inference Cloud Platform, building and operating the cloud layer behind an inference service. Tackle distributed systems challenges including multi-region traffic, active-active reliability, graceful degradation under bursty workloads, and high-QPS performance. Design core components like service discovery, routing, load balancing, caching, batching, and traffic management, while driving observability, incident response, capacity planning, and technical direction across adjacent teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 years ago

Staff Software Engineer, Inference Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own major architectural areas of the Inference Cloud Platform, building and operating the cloud layer behind an inference service. Tackle distributed systems challenges including multi-region traffic, active-active reliability, graceful degradation under bursty workloads, and high-QPS performance. Design core components like service discovery, routing, load balancing, caching, batching, and traffic management, while driving observability, incident response, capacity planning, and technical direction across adjacent teams.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Shape the technical direction and roadmap for the Inference Cloud Platform, including multi-region topology, failure domains, service boundaries, and system evolution.
  • •Design and build core platform components for AI inference workloads, including service discovery, request routing, load balancing, caching, batching, and traffic management.
  • •Architect reliability and performance for active-active systems with rapid failover, graceful degradation, and clear SLOs; improve latency, throughput, capacity efficiency, and resilience.
  • •Define traffic control and service-tier mechanisms such as admission control, quota management, rate limiting, and differentiated quality of service.
  • •Write and review production code on critical paths; lead hardest production issues with observability, incident response, capacity planning, and post-incident improvements, while mentoring senior engineers.

Key Requirements

  • •8+ years of software engineering experience with substantial individual contributor work building and operating large-scale distributed systems or cloud infrastructure.
  • •Deep expertise in distributed systems architecture in cloud environments, including networking, compute orchestration, container platforms, and multi-region production services.
  • •Proven ability to make sound architectural decisions for highly available, latency-sensitive systems at scale.
  • •Experience optimizing latency, throughput, and efficiency in high-QPS systems, including TTFT and tail-latency reduction.
  • •Strong proficiency in backend or systems languages such as Go, C++, or Python, contributing production code directly.
Experience:Distributed systemsCloud infrastructureAI inferenceMachine learning
Skills:Architectural judgmentTechnical leadershipDebuggingCommunicationMentorship
Tech Stack:GoC++PythonService discoveryRequest routingLoad balancingCachingBatchingTraffic managementActive-active systemsNetworkingCompute orchestrationContainer platformsObservabilityMetricsLoggingTracingAlertingSLOsRate limiting

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn