AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Independent problem solving","Engineering judgment","Debugging","Problem-solving","Root-cause analysis"]

Build and operate the platform layer that powers engineering infrastructure, including CI/CD systems and Kubernetes-based services. Create deployment automation and self-service workflows that make infrastructure changes repeatable, reviewable, and safe. Improve reliability, capacity, performance, cost efficiency, monitoring, and operational readiness while debugging issues across CI pipelines, Kubernetes, networking, storage, authentication, OS, and distributed applications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 month ago

AI Inference Core - Senior SW Engineer for Platform & DevOps

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and operate the platform layer that powers engineering infrastructure, including CI/CD systems and Kubernetes-based services. Create deployment automation and self-service workflows that make infrastructure changes repeatable, reviewable, and safe. Improve reliability, capacity, performance, cost efficiency, monitoring, and operational readiness while debugging issues across CI pipelines, Kubernetes, networking, storage, authentication, OS, and distributed applications.
Location: Sunnyvale, Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain CI/CD systems for build, test, integration, qualification, and release workflows.
  • •Build and operate Kubernetes-based platforms and services used by engineering teams across Cerebras.
  • •Develop deployment systems, internal tools, and self-service workflows for safe, repeatable infrastructure changes.
  • •Improve infrastructure reliability, capacity, performance, cost efficiency, monitoring, and operational readiness.
  • •Debug issues across CI pipelines, Kubernetes workloads, networking, storage, authentication, operating systems, and distributed applications, and implement lasting fixes.

Key Requirements

  • •3+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering.
  • •Hands-on experience building and maintaining CI/CD pipelines and automated software-delivery workflows.
  • •Experience deploying and operating services using Kubernetes and containerized environments.
  • •Experience with a major cloud platform, preferably AWS, and programmatic infrastructure provisioning.
  • •Strong understanding of Linux/Unix fundamentals and networking concepts (DNS, routing, load balancing, proxies, ports, TLS, and service connectivity).
Experience:3+ yearsPlatform engineeringDevOpsInfrastructure engineeringSite reliability engineeringSoftware engineering
Skills:Independent problem solvingEngineering judgmentDebuggingProblem-solvingRoot-cause analysis
Tech Stack:PythonShellCI/CDKubernetesContainerized environmentsAWSLinuxUnixDNSRoutingLoad balancingProxiesTLSLinux/Unix operating system fundamentalsMonitoringLoggingAlertingDashboardsIncident investigationTerraform

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn