Sr. Member of Technical Staff

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 1+ yearsEducation: mastersSkills: ["Collaboration","Troubleshooting","Documentation","Automation mindset"]

Design and build software features that improve system resiliency and high availability for distributed AI inference services. Develop AWS-based deployment workflows, Python scripts and APIs for real-time inference preprocessing/execution/post-processing, and automation to detect and mitigate failure modes. Work with Docker, Kubernetes, and monitoring/observability tooling to ensure scalable, low-latency performance, while debugging deployment and networking issues and documenting workflows, APIs, and release notes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
4 months ago

Sr. Member of Technical Staff

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design and build software features that improve system resiliency and high availability for distributed AI inference services. Develop AWS-based deployment workflows, Python scripts and APIs for real-time inference preprocessing/execution/post-processing, and automation to detect and mitigate failure modes. Work with Docker, Kubernetes, and monitoring/observability tooling to ensure scalable, low-latency performance, while debugging deployment and networking issues and documenting workflows, APIs, and release notes.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and develop features for system resiliency and high availability, including automated recovery mechanisms and fault-tolerant distributed architectures.
  • •Build and maintain AWS-based cloud deployment workflows for low-latency, scalable AI inference performance.
  • •Develop Python scripts and APIs for real-time inference data preprocessing, inference execution, and post-processing.
  • •Use parallel programming (multi-threading and asynchronous processing) to maximize resource efficiency on AWS compute instances.
  • •Develop and operate inference services using Docker and Kubernetes, including orchestration strategies, debugging, and defect triage using logs/metrics/traces.

Key Requirements

  • •Master’s degree (or foreign equivalent) in Computer Science or a related field.
  • •18 months of experience as an Information Security Analyst, Software Engineer, Sr. Member of Technical Staff, IT Senior Applications Engineer, or related occupation.
  • •Infrastructure-as-Code and deployment automation experience with Terraform, AWS CloudFormation, AWS CDK, and Ansible.
  • •Containerization and orchestration experience with Docker, Kubernetes, AWS EKS/ECS, AWS Fargate, and Helm.
  • •Programming and observability skills including Python, Node.js/JavaScript/Flask, and monitoring/logging/distributed tracing with CloudWatch, X-Ray, ELK, Prometheus, and Grafana.
Experience:1+ yearsAIInferenceCloudDistributed systems
Education:Master's in Computer Science (or a related field)
Skills:CollaborationTroubleshootingDocumentationAutomation mindset
Tech Stack:AWSTerraformAWS CloudFormationAWS CDKAnsibleDockerKubernetesAWS EKSAWS Elastic Container Service (ECS)AWS FargateHelmAWS EC2AWS LambdaAuto Scaling GroupsAWS CloudWatchAWS X-RayELK (Elasticsearch, Logstash, Kibana)PrometheusGrafanaPython

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn