Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI
San Francisco
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Root cause analysis","Communication","Prioritization"]

Own the reliability and operation of GPU cluster compute environments powering foundation model training and research. Debug issues across compute, storage, networking, schedulers, and distributed workloads, while improving CPU/GPU/storage utilization through tooling and automation. Build monitoring and validation, onboard and migrate workloads across GPU providers and hardware platforms, and contribute to the longer-term training infrastructure architecture.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Liquid AI
Liquid AI
1 month ago

Member of Technical Staff - GPU Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Own the reliability and operation of GPU cluster compute environments powering foundation model training and research. Debug issues across compute, storage, networking, schedulers, and distributed workloads, while improving CPU/GPU/storage utilization through tooling and automation. Build monitoring and validation, onboard and migrate workloads across GPU providers and hardware platforms, and contribute to the longer-term training infrastructure architecture.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own the reliability and operation of GPU clusters used for training and research.
  • •Debug issues across compute, storage, networking, schedulers, and distributed workloads.
  • •Improve CPU/GPU/storage utilization via better tooling and automation.
  • •Onboard and migrate workloads across GPU providers and hardware platforms.
  • •Build monitoring, validation, and platform abstractions and contribute to long-term training infrastructure architecture.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kPaid Leave

Key Requirements

  • •Strong software engineering experience building production-quality infrastructure tooling and automation.
  • •Deep knowledge of distributed systems, Linux, networking, and storage.
  • •Experience operating a shared compute cluster or distributed training platform.
  • •Proven ability to support production users and convert recurring failures into durable solutions.
  • •Technical depth to partner effectively with senior research and infrastructure engineers.
Experience:Distributed systemsAI training infrastructureHPCGPU infrastructure
Skills:Root cause analysisCommunicationPrioritization
Tech Stack:LinuxNetworkingStorageSLURMKubernetesRayHadoopDistributed systemsGPUHPC

Company Brief

Liquid AI
Builds AI infrastructure and tooling to enable real-time, distributed machine learning and orchestration across edge and cloud environments, simplifying deployment and management of intelligent applications.
Industry: AI & Machine Learning
Website