AI Research Engineer (Model Compression & Quantization)

Tether.io
Barcelona
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["PyTorch","Quantization","Distillation","Pruning","Transformers","C++"]

Join an AI research team focused on model compression and efficient deployment for multimodal AI systems (LLMs/VLMs). You will develop and evaluate quantization, distillation, and pruning strategies to reduce model footprint and latency on edge devices, build robust compression pipelines, and publish findings while advancing state-of-the-art in fintech AI applications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join an AI research team focused on model compression and efficient deployment for multimodal AI systems (LLMs/VLMs). You will develop and evaluate quantization, distillation, and pruning strategies to reduce model footprint and latency on edge devices, build robust compression pipelines, and publish findings while advancing state-of-the-art in fintech AI applications.
Location: Barcelona
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
  • •Experience with PyTorch deep learning frameworks or equivalent frameworks
  • •Hands-on experience with model quantization including both Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ).
  • •Research and hands-on experience with knowledge distillation for compressing large models into smaller, efficient ones.
  • •Research and hands-on experience with model pruning for compressing large models into smaller, efficient ones.
Experience:FintechMultimodalAI
Education:Bachelor's
Skills:PyTorchQuantizationDistillationPruningTransformersC++
Languages:English
Tech Stack:PyTorchQuantizationDistillationPruningTransformersC++

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website