AI Research Engineer (Model Compression & Quantization)

Tether.io
Abu Dhabi
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["PyTorch","Communication","Problem-solving","Team collaboration"]

Join an AI research team focused on reducing model size and compute for multimodal AI systems (LLMs/VLMs) through quantization, distillation, and pruning. You will build robust compression pipelines to run high-fidelity models on edge devices, balance latency and accuracy, and publish findings at top conferences.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join an AI research team focused on reducing model size and compute for multimodal AI systems (LLMs/VLMs) through quantization, distillation, and pruning. You will build robust compression pipelines to run high-fidelity models on edge devices, balance latency and accuracy, and publish findings at top conferences.
Location: Abu Dhabi
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field; ideally PhD in NLP, Machine Learning, or related field with publications in top conferences
Experience:MultimodalLLMsVLMsAI R&DEdge devices
Education:Bachelor's
Skills:PyTorchCommunicationProblem-solvingTeam collaboration
Languages:English
Tech Stack:PyTorchQATPTQQuantizationKnowledge distillationPruningMixed-precisionTransformersLLMsVLMsEdge devicesC++

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website