AI Research Engineer (Model Compression & Quantization)

Tether.io
Rome
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Problem-solving","Collaboration","Research"]

Join a globally distributed AI research team focused on model compression for multimodal systems. Lead quantization, distillation, and pruning efforts to reduce model size and latency on edge devices while preserving accuracy; publish findings and collaborate across R&D to advance state-of-the-art in LLM/VLM compression and deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
4 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Join a globally distributed AI research team focused on model compression for multimodal systems. Lead quantization, distillation, and pruning efforts to reduce model size and latency on edge devices while preserving accuracy; publish findings and collaborate across R&D to advance state-of-the-art in LLM/VLM compression and deployment.
Location: Rome
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field; ideally PhD in NLP, ML, or related field with publications in top venues
  • •Experience with PyTorch or equivalent frameworks
  • •Hands-on experience with model quantization (QAT and PTQ)
  • •Research and hands-on experience with knowledge distillation for compressing large models
  • •Research and hands-on experience with model pruning for compressing large models
Experience:FintechAIMultimodalEdge devices
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaborationResearch
Languages:English
Tech Stack:PythonPyTorchQATPTQQuantizationKnowledge distillationPruningTransformersC++

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website