AI Research Engineer (Kernel & Inference Optimization)
Tether.io
Bengaluru
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["GPU","Kernel optimization","Model serving","Inference","MSL","Quantization","Pruning","Diffusion models","Vision transformers","Flash attention","KV cache","Eagle"]Design and optimize model serving and inference pipelines for edge/mobile AI deployment across resource-constrained devices, from kernel-level optimizations to end-to-end serving architectures. You will implement low-latency, high-throughput inference on diverse hardware, collaborate with cross-functional teams, and advance state-of-the-art techniques for scalable AI on mobile platforms, handling multi-modal data and evaluating performance with real-world metrics.

