Training / AI Infrastructure

Genesis AI
United States
Workplace: OnsiteFull timeFunction: Education & TrainingExperience: 8+ yearsSkills: ["Profiling","Performance optimization","Debugging","Systems thinking"]

Build and optimize distributed foundation model training systems to reduce wall-clock convergence time. You’ll design multi-node PyTorch training for scalability and high GPU utilization, implement performance-critical CUDA/cuDNN/Triton and custom kernels, and tune CPU/GPU workloads, memory, throughput, and networking. You’ll also develop monitoring and debugging tools to quickly diagnose performance regressions and failures across large-scale runs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Genesis AI
Genesis AI
4 months ago

Training / AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and optimize distributed foundation model training systems to reduce wall-clock convergence time. You’ll design multi-node PyTorch training for scalability and high GPU utilization, implement performance-critical CUDA/cuDNN/Triton and custom kernels, and tune CPU/GPU workloads, memory, throughput, and networking. You’ll also develop monitoring and debugging tools to quickly diagnose performance regressions and failures across large-scale runs.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Profile and eliminate bottlenecks across the foundation model training stack, from data pipelines to GPU kernels.
  • •Design, build, and optimize distributed training systems for multi-node GPU clusters.
  • •Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it into training frameworks.
  • •Optimize workloads for hardware efficiency, including CPU/GPU compute balance, memory management, data throughput, and networking.
  • •Develop monitoring and debugging tools to diagnose performance regressions and failures in large-scale runs.

Key Requirements

  • •Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years).
  • •Production-grade expertise in Python.
  • •Low-level performance mastery including CUDA/cuDNN/Triton, CPU–GPU interactions, data movement, and kernel optimization.
  • •Experience scaling frontier training jobs using PyTorch and data, context, pipeline, and model parallelism.
  • •A system-level mindset with a track record of tuning hardware–software interactions for maximum utilization.
Experience:8+ yearsMachine learning infrastructureHigh-performance computingDistributed systemsPyTorch
Skills:ProfilingPerformance optimizationDebuggingSystems thinking
Tech Stack:PythonPyTorchCUDACuDNNTritonCustom kernelsGPU

Company Brief

Genesis AI
Genesis AI is a global physical AI research lab and full‑stack robotics company building a universal robotics foundation model and horizontal platform to enable general‑purpose robots and scale automation of physical labor.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2024
WebsiteLinkedIn