AI Research Engineer (Model Compression & Quantization)

Tether.io
Zurich, Belgium
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Collaboration","Problem-solving"]

Join a distributed AI research team focusing on model compression for multimodal systems. You will develop and deploy quantization, distillation, and pruning techniques to reduce model size and latency for LLMs and VLMs, balance accuracy with efficiency, publish findings, and drive production-ready edge deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Join a distributed AI research team focusing on model compression for multimodal systems. You will develop and deploy quantization, distillation, and pruning techniques to reduce model size and latency for LLMs and VLMs, balance accuracy with efficiency, publish findings, and drive production-ready edge deployments.
Location: Zurich, Belgium
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field; ideally a PhD in NLP, Machine Learning, or related field with strong AI R&D publications.
  • •Experience with PyTorch or equivalent DL frameworks.
  • •Hands-on experience with model quantization including Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ).
  • •Research and hands-on experience with knowledge distillation for compressing large models into smaller, efficient ones.
  • •Research and hands-on experience with model pruning for compressing large models into smaller, efficient ones.
Experience:MultimodalAINLPDeep learning
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PyTorchQuantizationQATPTQKnowledge distillationPruningTransformersC++

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website