Inference

Genesis AI
United States
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 8+ yearsSkills: ["System-level thinking","Performance optimization","Monitoring and debugging","Reliability focus","Regression troubleshooting"]

Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to improve throughput with large-batch serving. Implement performance-critical code using CUDA, Triton, and custom kernels, then tune workloads for both latency and throughput. Develop monitoring and debugging tools to ensure reliability, determinism, and fast regression diagnosis across the full stack.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Genesis AI
Genesis AI
4 months ago

Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to improve throughput with large-batch serving. Implement performance-critical code using CUDA, Triton, and custom kernels, then tune workloads for both latency and throughput. Develop monitoring and debugging tools to ensure reliability, determinism, and fast regression diagnosis across the full stack.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build low-latency inference pipelines for on-device deployment in robotics.
  • •Design and optimize distributed inference systems on GPU clusters to improve throughput.
  • •Implement efficient low-level code with CUDA, Triton, and custom kernels and integrate with high-level frameworks.
  • •Optimize inference workloads for both throughput and latency (batching, scheduling, quantization, caching, memory management, graph compilation).
  • •Develop monitoring and debugging tools to ensure reliability, determinism, and rapid diagnosis of regressions across stacks.

Key Requirements

  • •Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years).
  • •Production-grade Python with strong systems-language background (C++/Rust/Go).
  • •Low-level performance expertise including CUDA, Triton, kernel optimization, quantization, and memory/compute scheduling.
  • •Track record scaling inference workloads for both throughput-oriented clusters and latency-critical on-device deployment.
  • •System-level mindset for tuning hardware–software interactions to maximize efficiency and responsiveness.
Experience:8+ yearsDistributed systemsML infrastructureHigh-performance servingRoboticsGPU clustersOn-device inference
Skills:System-level thinkingPerformance optimizationMonitoring and debuggingReliability focusRegression troubleshooting
Tech Stack:PythonC++RustGoCUDATritonCustom kernelsQuantizationGPU clustersDistributed systemsNext-tokenDiffusion-based control loopsGraph compilationCaching

Company Brief

Genesis AI
Genesis AI is a global physical AI research lab and full‑stack robotics company building a universal robotics foundation model and horizontal platform to enable general‑purpose robots and scale automation of physical labor.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2024
WebsiteLinkedIn