AI Research Engineer (Model Compression & Quantization)

Tether.io
Rome, London
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Problem-solving","Research","Collaboration"]

Join Tether's AI research team to drive model compression for multimodal AI, focusing on quantization, knowledge distillation, and pruning to run large language models and vision-language models efficiently on edge devices. Build robust compression pipelines, establish performance metrics, and balance model size, latency, and accuracy while publishing impactful results. This is a remote-friendly role with a global team shaping fintech AI.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Join Tether's AI research team to drive model compression for multimodal AI, focusing on quantization, knowledge distillation, and pruning to run large language models and vision-language models efficiently on edge devices. Build robust compression pipelines, establish performance metrics, and balance model size, latency, and accuracy while publishing impactful results. This is a remote-friendly role with a global team shaping fintech AI.
Location: Rome, London
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or related field with publications in top conferences.
  • •Experience with PyTorch or equivalent frameworks
  • •Hands-on experience with model quantization including Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ)
  • •Research and hands-on experience with knowledge distillation for compressing large models
  • •Research and hands-on experience with model pruning for compressing large models
Experience:MultimodalAIEdge devices
Education:Bachelor's
Skills:Problem-solvingResearchCollaboration
Languages:English
Tech Stack:PyTorchQuantizationKnowledge DistillationPruningTransformersC++

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website