Training / AI Infrastructure
London
Workplace: HybridFull timeFunction: Education & TrainingExperience: 8+ yearsSkills: []Build and optimize distributed training infrastructure for frontier AI models, reducing wall-clock time to convergence. You’ll profile and eliminate bottlenecks across the training stack—from data pipelines to GPU kernels—while designing multi-node PyTorch systems that scale efficiently. Develop low-level CUDA/cuDNN/Triton/custom-kernel performance improvements and integrate them into high-level training frameworks. Create monitoring and debugging tools to quickly diagnose performance regressions and failures.
Loading
Loading job details...
Preparing the role view and application actions.

