Member of Technical Staff - Engineering Lead, Compute Platform

Reflection AI
San Francisco, London
Workplace: OnsiteFull timeFunction: Administration & Executive AssistanceSkills: ["Hiring","Mentoring","Technical leadership","Strategic planning","Cross-team collaboration"]

Own the compute platform layer end to end, setting technical direction for a Kubernetes-based, multi-cloud environment that powers large-scale training. Lead and mentor systems engineers, define architecture and multi-year strategy, and drive cluster management tools for remediation, topology-aware scheduling, capacity planning, and hardware debugging. Partner with training teams on fault tolerance, node health checks, and remediation while ensuring robust monitoring and roadmap execution for next-gen GPUs and storage replication.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
1 month ago

Member of Technical Staff - Engineering Lead, Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Own the compute platform layer end to end, setting technical direction for a Kubernetes-based, multi-cloud environment that powers large-scale training. Lead and mentor systems engineers, define architecture and multi-year strategy, and drive cluster management tools for remediation, topology-aware scheduling, capacity planning, and hardware debugging. Partner with training teams on fault tolerance, node health checks, and remediation while ensuring robust monitoring and roadmap execution for next-gen GPUs and storage replication.
Location: San Francisco, London
Workplace: Onsite
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Lead the Compute Platform team by hiring, mentoring, and setting technical standards, planning, and prioritization.
  • •Own the technical direction and multi-year architecture for compute platform reliability, multi-cloud scheduling, and cluster management.
  • •Drive tools and processes for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging.
  • •Guide platform engineering for cluster management across large, multi-GPU fleets and workloads.
  • •Ensure comprehensive cluster-wide monitoring and partner with training teams on fault tolerance, node health checks, and remediation strategies.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityHealth InsuranceDentalVisionLifePaid ParentalPaid Leave

Key Requirements

  • •Track record leading systems or infrastructure teams (hiring, mentoring, and setting technical direction) while remaining hands-on.
  • •Deep systems-level engineering experience focused on cluster-wide behavior, maintenance, and reliability.
  • •Strong coding ability with demonstrated focus on systems or GPU infrastructure.
  • •Deep GPU hardware knowledge beyond standard Kubernetes, including familiarity with NCCL.
  • •Ability to set strategy and drive execution across a multi-cloud, large-fleet environment, partnering effectively with research and training teams.
Skills:HiringMentoringTechnical leadershipStrategic planningCross-team collaboration
Tech Stack:KubernetesK8sMulti-cloud schedulingNCCLCluster managementGPU-to-GPU networkMulti-GPU fleetsCloud storageVAST

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor