Training / AI Infrastructure

Genesis AI
Paris, London
Workplace: HybridFull timeFunction: Education & TrainingExperience: 8+ yearsSkills: ["Problem-solving","Analytical thinking","Performance optimization","Debugging"]

Work on foundation model training performance by profiling and removing bottlenecks across the training stack, from data pipelines to GPU kernels. Build and optimize distributed training systems for multi-node GPU clusters using PyTorch, and implement efficient low-level components with CUDA, cuDNN, Triton, and custom kernels. Improve hardware utilization through workload and memory optimizations, and create monitoring/debugging tools to quickly diagnose performance regressions and failures.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Genesis AI
Genesis AI
4 months ago

Training / AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Work on foundation model training performance by profiling and removing bottlenecks across the training stack, from data pipelines to GPU kernels. Build and optimize distributed training systems for multi-node GPU clusters using PyTorch, and implement efficient low-level components with CUDA, cuDNN, Triton, and custom kernels. Improve hardware utilization through workload and memory optimizations, and create monitoring/debugging tools to quickly diagnose performance regressions and failures.
Location: Paris, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Profile and eliminate bottlenecks across the foundation model training stack, from data pipelines to GPU kernels.
  • •Design, build, and optimize distributed training systems in PyTorch for multi-node GPU clusters.
  • •Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it into high-level training frameworks.
  • •Optimize workloads for hardware efficiency, including CPU/GPU compute balance, memory management, data throughput, and networking.
  • •Develop monitoring and debugging tools for large-scale runs to diagnose performance regressions and failures quickly.

Key Requirements

  • •8+ years of experience in distributed systems, ML infrastructure, or high-performance computing.
  • •Production-grade expertise in Python.
  • •Strong low-level performance skills with CUDA/cuDNN/Triton, including CPU–GPU interactions and kernel optimization.
  • •Experience scaling frontier training jobs with PyTorch (data, context, pipeline, and model parallelism).
  • •Demonstrated system-level mindset for tuning hardware–software interactions to maximize utilization.
Experience:8+ yearsMachine learning infrastructureDistributed systemsHigh-performance computingFoundation model training
Skills:Problem-solvingAnalytical thinkingPerformance optimizationDebugging
Tech Stack:PythonPyTorchCUDACuDNNTritonGPU kernelsDistributed trainingMulti-node GPU clustersCPU–GPU interactionsCustom kernelsCUDA kernels

Company Brief

Genesis AI
Genesis AI is a global physical AI research lab and full‑stack robotics company building a universal robotics foundation model and horizontal platform to enable general‑purpose robots and scale automation of physical labor.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2024
WebsiteLinkedIn