AI Research Engineer (Model Compression & Quantization)

Tether.io
Dubai
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Problem-solving"]

Design and implement model compression techniques for multimodal AI (LLMs, VLMs) at Tether, focusing on quantization, knowledge distillation, and pruning to reduce model size and latency. Build robust pipelines, measure performance vs. accuracy, and publish findings while collaborating with a global remote team to advance edge deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
4 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Design and implement model compression techniques for multimodal AI (LLMs, VLMs) at Tether, focusing on quantization, knowledge distillation, and pruning to reduce model size and latency. Build robust pipelines, measure performance vs. accuracy, and publish findings while collaborating with a global remote team to advance edge deployments.
Location: Dubai
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or related field, with a solid track record in AI R&D (publications in top conferences).
  • •Experience with PyTorch deep learning frameworks or equivalent frameworks
  • •Hands-on experience with model quantization including both Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ).
  • •Research and hands-on experience with knowledge distillation for compressing large models into smaller, efficient ones.
  • •Research and hands-on experience with model pruning for compressing large models into smaller, efficient ones.
Experience:AIMultimodalLLMsVLMsModel compression
Education:Bachelor's
Skills:CommunicationProblem-solving
Languages:English
Tech Stack:PyTorchTransformersC++QATPTQDistillationPruning

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website