Member of Technical Staff - Compute Platform

Reflection AI
New York, London, San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Collaboration","Problem-solving"]

Join the Compute Platform team to keep a Kubernetes-based multi-cloud compute layer healthy and highly available for large-scale GPU workloads. You’ll build tooling for automatic remediation, design cluster-management Stack improvements, implement monitoring and observability, and roadmap GPU deployments and cross-data-center storage. You’ll collaborate with training teams to optimize fault tolerance and performance across multi-GPU fleets.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
4 months ago

Member of Technical Staff - Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Join the Compute Platform team to keep a Kubernetes-based multi-cloud compute layer healthy and highly available for large-scale GPU workloads. You’ll build tooling for automatic remediation, design cluster-management Stack improvements, implement monitoring and observability, and roadmap GPU deployments and cross-data-center storage. You’ll collaborate with training teams to optimize fault tolerance and performance across multi-GPU fleets.
Location: New York, London, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Cluster Management: Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning and rapid hardware debugging.
  • •Platform Engineering: Design and iterate on our cluster management stack for workloads across large, multi-GPU fleets.
  • •Monitoring & Observability: Implement comprehensive cluster-wide monitoring, focusing on durability and active performance benchmarking.
  • •Roadmap Execution: Prepare the infrastructure for next-generation GPU deployments and increasingly larger cluster sizes; own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.
  • •Collaborate with training teams to co-design fault tolerance, node health checks, and remediation strategies.

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsuranceDisability InsuranceParental LeaveRelocationPaid Leave

Key Requirements

  • •Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
  • •Strong coding ability with a demonstrated focus on systems or GPU infrastructure.
  • •Deep GPU hardware knowledge beyond standard Kubernetes, e.g., NCCL.
  • •Alignment with a Kubernetes-first architecture.
  • •Cloud storage expertise, specifically managing high-performance data products across multiple data centers, connecting storage environments and handling datasets and checkpointing at scale.
Experience:AICloudDistributed Systems
Skills:CollaborationProblem-solving
Languages:English
Tech Stack:KubernetesNCCLGPUMulti-cloudStorageVASTCheckpointing

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor