AI Research Engineer (Model Compression & Quantization)

Tether.io
Brazil
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Collaboration","Research"]

Research and develop model compression techniques for multimodal AI (LLMs and VLMs) to reduce footprint and latency. Build robust compression pipelines using quantization, knowledge distillation, and pruning; evaluate trade-offs between size, speed, and accuracy; stay at the cutting edge with new techniques and publish findings. Remote role with a global, async team and edge deployment focus.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Research and develop model compression techniques for multimodal AI (LLMs and VLMs) to reduce footprint and latency. Build robust compression pipelines using quantization, knowledge distillation, and pruning; evaluate trade-offs between size, speed, and accuracy; stay at the cutting edge with new techniques and publish findings. Remote role with a global, async team and edge deployment focus.
Location: Brazil
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models for efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies to optimize the accuracy–performance balance.

Key Requirements

  • •A degree in Computer Science or related field, ideally PhD in NLP/Machine Learning with publications in top conferences
  • •Experience with PyTorch or equivalent deep learning frameworks
  • •Hands-on experience with model quantization (QAT and PTQ)
  • •Experience with knowledge distillation for compressing large models
  • •Experience with model pruning to reduce parameters and computation
Experience:FintechAI
Education:Bachelor's
Skills:CommunicationCollaborationResearch
Tech Stack:PyTorchC++QuantizationQATPTQ

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website