AI Research Engineer (Model Compression & Quantization)

Tether.io
Dublin
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["PyTorch","Quantization","Quantization-Aware Training","QAT","Post-Training Quantization","PTQ","Knowledge distillation","Pruning","Transformers","NLP","Machine learning","Fine-tuning","C++"]

Join an AI research team focused on model compression and efficient deployment for multimodal AI systems (LLMs, VLMs) to reduce footprint and latency while preserving accuracy. You will develop quantization, distillation, and pruning pipelines, experiment with mixed-precision strategies, publish results, and contribute to cutting-edge compression techniques for edge devices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Join an AI research team focused on model compression and efficient deployment for multimodal AI systems (LLMs, VLMs) to reduce footprint and latency while preserving accuracy. You will develop quantization, distillation, and pruning pipelines, experiment with mixed-precision strategies, publish results, and contribute to cutting-edge compression techniques for edge devices.
Location: Dublin
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field; ideally PhD in NLP, ML, or related field with strong AI R&D publications.
Experience:Multimodal AIEdge AITransformers
Education:Bachelor's
Skills:PyTorchQuantizationQuantization-Aware TrainingQATPost-Training QuantizationPTQKnowledge distillationPruningTransformersNLPMachine learningFine-tuningC++
Languages:English
Tech Stack:PythonPyTorchC++TransformersQATPTQDistillationPruning

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website