ML Infra Engineer

Physical Intelligence
San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Communication","Ownership","Problem-solving"]

You will design, build, and maintain large-scale ML training infrastructure, owning systems for training and inference, scaling JAX/PyTorch training across TPU/GPU clusters, and optimizing performance. You’ll collaborate with researchers and engineers to turn ideas into experiments and production runs, balancing researcher flexibility with system reliability in a fast-paced ML infrastructure team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Physical Intelligence
Physical Intelligence
2 years ago

ML Infra Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

You will design, build, and maintain large-scale ML training infrastructure, owning systems for training and inference, scaling JAX/PyTorch training across TPU/GPU clusters, and optimizing performance. You’ll collaborate with researchers and engineers to turn ideas into experiments and production runs, balancing researcher flexibility with system reliability in a fast-paced ML infrastructure team.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.
  • •Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction.
  • •Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization.
  • •Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • •Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost.

Key Requirements

  • •Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
  • •Hands-on large-scale training experience in JAX (preferred) or PyTorch.
  • •Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines.
  • •Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS).
  • •Ability to debug and optimize performance bottlenecks across the training stack.
Experience:Machine learningInfrastructureCloud
Skills:CommunicationOwnershipProblem-solving
Tech Stack:JAXPyTorchTPUGPUSLURMKubernetesGCPAWS

Company Brief

Physical Intelligence
Develops large-scale AI models and learning algorithms to enable general-purpose intelligence for robots and physically-actuated devices, aiming to power a wide range of real-world robotic tasks and applications.
Industry: Robotics
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2024
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn