Training / AI Infrastructure

Genesis AI
London
Workplace: HybridFull timeFunction: Education & TrainingExperience: 8+ yearsSkills: []

Build and optimize distributed training infrastructure for frontier AI models, reducing wall-clock time to convergence. You’ll profile and eliminate bottlenecks across the training stack—from data pipelines to GPU kernels—while designing multi-node PyTorch systems that scale efficiently. Develop low-level CUDA/cuDNN/Triton/custom-kernel performance improvements and integrate them into high-level training frameworks. Create monitoring and debugging tools to quickly diagnose performance regressions and failures.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Genesis AI
Genesis AI
2 days ago

Training / AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and optimize distributed training infrastructure for frontier AI models, reducing wall-clock time to convergence. You’ll profile and eliminate bottlenecks across the training stack—from data pipelines to GPU kernels—while designing multi-node PyTorch systems that scale efficiently. Develop low-level CUDA/cuDNN/Triton/custom-kernel performance improvements and integrate them into high-level training frameworks. Create monitoring and debugging tools to quickly diagnose performance regressions and failures.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Profile and eliminate bottlenecks across the foundation model training stack, from data pipelines to GPU kernels, to reduce time to convergence.
  • •Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization.
  • •Implement and optimize low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it into high-level training frameworks.
  • •Optimize workloads for hardware efficiency across CPU/GPU balance, memory management, data throughput, and networking.
  • •Develop monitoring and debugging tools to enable rapid diagnosis of performance regressions and training failures.

Key Requirements

  • •Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years).
  • •Production-grade expertise in Python.
  • •Low-level performance mastery across CUDA/cuDNN/Triton, including CPU–GPU interactions, data movement, and kernel optimization.
  • •Experience scaling training jobs with PyTorch, including data, context, pipeline, and model parallelism.
  • •A system-level mindset with a proven track record of tuning hardware–software interactions for maximum utilization.
Experience:8+ yearsDistributed systemsML infrastructureHigh-performance computing
Tech Stack:PythonPyTorchCUDACuDNNTriton

Company Brief

Genesis AI
Genesis AI is a global physical AI research lab and full‑stack robotics company building a universal robotics foundation model and horizontal platform to enable general‑purpose robots and scale automation of physical labor.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2024
WebsiteLinkedIn