AI Research Engineer (Model Compression & Quantization)

Tether.io
Amsterdam
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: high_schoolSkills: ["Communication","Collaboration","Problem-solving"]

Join an AI research team focused on model compression and efficient deployment for multimodal systems (LLMs, VLMs). You will develop quantization, distillation, and pruning pipelines to reduce footprint and latency on edge devices while preserving accuracy. Collaborate with global teams to publish findings and advance state-of-the-art in scalable AI for fintech applications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join an AI research team focused on model compression and efficient deployment for multimodal systems (LLMs, VLMs). You will develop quantization, distillation, and pruning pipelines to reduce footprint and latency on edge devices while preserving accuracy. Collaborate with global teams to publish findings and advance state-of-the-art in scalable AI for fintech applications.
Location: Amsterdam
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in top conferences).
  • •Experience with PyTorch deep learning frameworks or equivalent frameworks
  • •Hands-on experience with model quantization including both Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ).
  • •Research and hands-on experience with knowledge distillation for compressing large models into smaller, efficient ones.
  • •Research and hands-on experience with model pruning for compressing large models into smaller, efficient ones.
Experience:FintechAIMultimodalEdge devices
Education:High School
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonPyTorchC++QuantizationQATPTQKnowledge DistillationPruningTransformersLLMsVLMsMultimodalEdge devices

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website