AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether.io
Italy, Argentina, Colombia, Brazil, Uruguay, India, Bangladesh, Pakistan, Vietnam, Taiwan, Thailand, United Kingdom, Switzerland, Spain, United Arab Emirates, Bulgaria, Czech Republic, Serbia, Denmark, Estonia, Greece, Georgia, Budapest, Ireland, Malta, Norway, Amsterdam, Poland, Portugal, Romania, Sweden, Belgium, Israel, Cyprus
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Metal Shading Language","GPU kernels","Model serving","Inference optimization","Edge deployment","Latency optimization"]

Join an AI model team focusing on design and optimization of model serving, inference pipelines, and edge/on-device deployment. You will architect high-throughput, low-latency serving solutions, run controlled tests, evaluate memory usage, and collaborate with cross-functional teams to push the boundaries of model compression, quantization, and efficient AI across diverse hardware.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 45 minutes agoStatus: Live

Job Summary

Join an AI model team focusing on design and optimization of model serving, inference pipelines, and edge/on-device deployment. You will architect high-throughput, low-latency serving solutions, run controlled tests, evaluate memory usage, and collaborate with cross-functional teams to push the boundaries of model compression, quantization, and efficient AI across diverse hardware.
Location: Italy, Argentina, Colombia, Brazil, Uruguay, India, Bangladesh, Pakistan, Vietnam, Taiwan, Thailand, United Kingdom, Switzerland, Spain, United Arab Emirates, Bulgaria, Czech Republic, Serbia, Denmark, Estonia, Greece, Georgia, Budapest, Ireland, Malta, Norway, Amsterdam, Poland, Portugal, Romania, Sweden, Belgium, Israel, Cyprus
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage. Ensure these pipelines run efficiently across diverse environments, including resource-constrained devices and edge platforms.
  • •Build, run, and monitor controlled inference tests in both simulated and live production environments. Track key performance indicators such as response latency, throughput, memory consumption, and error rates, with special attention to metrics specific to resource-constrained devices.
  • •Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges, specifically those encountered on low-resource devices.
  • •Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring both processing and memory metrics. Address issues such as suboptimal batch processing, network delays, and high memory usage.
  • •Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines designed for edge and on-device applications. Define clear success metrics and ensure continuous monitoring and iterative refinements.

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
  • •Must have knowledge of Metal Shading Language (MSL). You should be comfortable writing custom compute shaders from scratch.
  • •Proven experience in low-level kernel optimizations and inference optimization on mobile devices. Your contributions should have led to measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications, particularly on resource-constrained devices and edge platforms.
  • •A deep understanding of modern model serving architectures and inference optimization techniques. This includes state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
  • •Must have strong expertise in writing GPU kernels for mobile devices (i.e., smartphones) as well as a deep understanding of model serving frameworks and engines. Practical experience in developing and deploying end-to-end inference pipelines, from optimizing models for efficient serving to integrating these solutions on resource-constrained devices is required.
Experience:AIModel compressionEdge computingMobile
Education:Bachelor's
Skills:Metal Shading LanguageGPU kernelsModel servingInference optimizationEdge deploymentLatency optimization
Languages:English
Tech Stack:Metal Shading LanguageGPU kernelsTensor ParallelismPipeline ParallelismVision TransformersDiffusion ModelsQuantizationPruningFlash attention

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website