ML Infra Engineer, Modeling

Physical Intelligence
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Cross-functional communication","Ownership mindset","Debugging","Performance optimization"]

Own and scale training and inference infrastructure for large-scale model development. Build and maintain reusable, efficient JAX training pipelines, orchestration, scheduling, checkpointing, and monitoring. Collaborate with researchers to scale distributed training across TPU/GPU clusters, improving performance through profiling and optimizations. Help evolve core training code so research experiments can reliably become production-grade training runs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Physical Intelligence
Physical Intelligence
5 days ago

ML Infra Engineer, Modeling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Own and scale training and inference infrastructure for large-scale model development. Build and maintain reusable, efficient JAX training pipelines, orchestration, scheduling, checkpointing, and monitoring. Collaborate with researchers to scale distributed training across TPU/GPU clusters, improving performance through profiling and optimizations. Help evolve core training code so research experiments can reliably become production-grade training runs.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own training and inference infrastructure, including scheduling, job management, checkpointing, and metrics/logging.
  • •Scale distributed training across TPU and GPU clusters with minimal friction.
  • •Profile and improve memory usage, device utilization, throughput, and distributed synchronization.
  • •Build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • •Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.

Key Requirements

  • •Strong software engineering fundamentals with experience building ML training infrastructure or internal platforms.
  • •Hands-on large-scale training experience with JAX (preferred) and/or PyTorch.
  • •Experience with distributed training (multi-host setups), data loaders, and evaluation pipelines.
  • •Ability to manage training workloads on cloud platforms such as SLURM, Kubernetes, GCP TPU/GKE, and/or AWS.
  • •Proven ability to debug and optimize performance bottlenecks across the training stack.
Experience:Machine learningDistributed trainingCloudRoboticsFoundation models
Skills:Cross-functional communicationOwnership mindsetDebuggingPerformance optimization
Tech Stack:JAXPyTorchTPUGPUSLURMKubernetesGCPGKEAWSData loaders

Company Brief

Physical Intelligence
Develops large-scale AI models and learning algorithms to enable general-purpose intelligence for robots and physically-actuated devices, aiming to power a wide range of real-world robotic tasks and applications.
Industry: Robotics
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2024
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn