Training / AI Infrastructure
Paris, London
Workplace: HybridFull timeFunction: Education & TrainingExperience: 8+ yearsSkills: ["Problem-solving","Analytical thinking","Performance optimization","Debugging"]Work on foundation model training performance by profiling and removing bottlenecks across the training stack, from data pipelines to GPU kernels. Build and optimize distributed training systems for multi-node GPU clusters using PyTorch, and implement efficient low-level components with CUDA, cuDNN, Triton, and custom kernels. Improve hardware utilization through workload and memory optimizations, and create monitoring/debugging tools to quickly diagnose performance regressions and failures.
Loading
Loading job details...
Preparing the role view and application actions.

