Platform Engineer

Zyphra
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["CI/CD","Infrastructure as code","Containers","Kubernetes","Docker","Terraform","Ansible","Slurm","GPU","VLLM","Ray","SGLang","Triton","Apptainer"]

Platform Engineer responsible for designing and maintaining Zyphra’s infrastructure to be robust, observable, secure, and scalable. Collaborates with ML, DevOps, and infra teams to ensure reliable ML workloads, secure release processes, and reproducible compute environments, with a focus on incident response and improving system performance for end users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zyphra
Zyphra
6 months ago

Platform Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Platform Engineer responsible for designing and maintaining Zyphra’s infrastructure to be robust, observable, secure, and scalable. Collaborates with ML, DevOps, and infra teams to ensure reliable ML workloads, secure release processes, and reproducible compute environments, with a focus on incident response and improving system performance for end users.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build and maintain observability systems including monitoring, logging, and alerting to ensure reliability of ML workloads
  • •Manage Infrastructure as a Service across the stack and CI/CD in collaboration with engineering teams
  • •Design resilient build and deployment systems across research and production environments
  • •Implement secure release processes with strong auditability and rollback support
  • •Lead incident response, root-cause analysis, and postmortems to drive learning and prevention

Pay and Benefits

Perks:Health InsuranceDentalVision401kPaid Leave

Key Requirements

  • •Experience in high-performance compute environments, such as ML clusters or GPU farms as well as hyperscaler cloud environments (i.e. AWS, GCP, etc.)
  • •Background in infrastructure as code (i.e., Terraform, Ansible, etc.)
  • •Familiarity with containers (i.e., Docker, Apptainer) and their integration with scheduling systems (i.e., Kubernetes, Slurm)
  • •Familiarity with software release engineering for ML/AI systems is a plus
  • •Experience managing run-books, DRP, change management, and general fault tolerance
Experience:Artificial intelligenceMl infrastructureCloud computing
Skills:CI/CDInfrastructure as codeContainersKubernetesDockerTerraformAnsibleSlurmGPUVLLMRaySGLangTritonApptainer
Languages:English
Tech Stack:TerraformAnsibleDockerApptainerKubernetesSlurmAWSGCPRaySGLangTritonVLLMCI/CDGitFSA

Company Brief

Zyphra
Zyphra is a technology company that provides software solutions and consulting services to help organizations implement digital tools and improve operational workflows. The company focuses on custom development, integration, and support for enterprise technology projects.
Industry: Consulting
Website