AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
Tether.io
Italy, Argentina, Colombia, Brazil, Uruguay, India, Bangladesh, Pakistan, Vietnam, Taiwan, Thailand, United Kingdom, Switzerland, Spain, United Arab Emirates, Bulgaria, Czech Republic, Serbia, Denmark, Estonia, Greece, Georgia, Budapest, Ireland, Malta, Norway, Amsterdam, Poland, Portugal, Romania, Sweden, Belgium, Israel, Cyprus
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Metal Shading Language","GPU kernels","Model serving","Inference optimization","Edge deployment","Latency optimization"]Join an AI model team focusing on design and optimization of model serving, inference pipelines, and edge/on-device deployment. You will architect high-throughput, low-latency serving solutions, run controlled tests, evaluate memory usage, and collaborate with cross-functional teams to push the boundaries of model compression, quantization, and efficient AI across diverse hardware.

