Inference

Genesis AI
Paris, London
Workplace: HybridFull timeFunction: Software EngineeringExperience: 8+ yearsSkills: ["System-level thinking","Performance optimization","Debugging","Reliability focus","Regression diagnosis"]

Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to increase throughput and efficiency using large-batch serving. Write and integrate low-level CUDA/Triton kernels, then tune workloads for both latency and throughput via batching, scheduling, quantization, caching, and graph compilation. Create monitoring and debugging tools to ensure reliability and rapid regression diagnosis.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Genesis AI
Genesis AI
4 months ago

Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to increase throughput and efficiency using large-batch serving. Write and integrate low-level CUDA/Triton kernels, then tune workloads for both latency and throughput via batching, scheduling, quantization, caching, and graph compilation. Create monitoring and debugging tools to ensure reliability and rapid regression diagnosis.
Location: Paris, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build low-latency inference pipelines for on-device deployment to support real-time next-token and diffusion-based robotics control loops.
  • •Design and optimize distributed inference systems on GPU clusters to improve throughput and resource utilization using large-batch serving.
  • •Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it into high-level frameworks.
  • •Optimize inference workloads for both throughput and latency using batching, scheduling, quantization, caching, memory management, and graph compilation.
  • •Develop monitoring and debugging tools to ensure reliability, determinism, and fast diagnosis of regressions across stacks.

Key Requirements

  • •Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years).
  • •Production-grade expertise in Python, with a strong background in systems languages (C++/Rust/Go).
  • •Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling.
  • •Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments.
  • •System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness.
Experience:8+ yearsDistributed systemsML infrastructureHigh-performance servingRoboticsOn-device deployments
Skills:System-level thinkingPerformance optimizationDebuggingReliability focusRegression diagnosis
Tech Stack:PythonC++RustGoCUDATritonCustom kernelsQuantizationGPU clustersGraph compilationCaching

Company Brief

Genesis AI
Genesis AI is a global physical AI research lab and full‑stack robotics company building a universal robotics foundation model and horizontal platform to enable general‑purpose robots and scale automation of physical labor.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2024
WebsiteLinkedIn