MTS, Training
Genesis AI
United States
Workplace: OnsiteFull timeFunction: Education & TrainingExperience: 8+ yearsSkills: ["Profiling","Performance optimization","Debugging","Systems thinking"]Build and optimize distributed foundation model training systems to reduce wall-clock convergence time. You’ll design multi-node PyTorch training for scalability and high GPU utilization, implement performance-critical CUDA/cuDNN/Triton and custom kernels, and tune CPU/GPU workloads, memory, throughput, and networking. You’ll also develop monitoring and debugging tools to quickly diagnose performance regressions and failures across large-scale runs.

