MTS, Inference
Genesis AI
United States
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 8+ yearsSkills: ["System-level thinking","Performance optimization","Monitoring and debugging","Reliability focus","Regression troubleshooting"]Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to improve throughput with large-batch serving. Implement performance-critical code using CUDA, Triton, and custom kernels, then tune workloads for both latency and throughput. Develop monitoring and debugging tools to ensure reliability, determinism, and fast regression diagnosis across the full stack.

