AI Research Engineer (Model Compression & Quantization)

Tether.io
Bengaluru
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Soft skills","Problem-solving"]

Join our AI research team to advance model compression for multimodal AI systems (LLMs and VLMs). You’ll design and implement quantization, distillation, and pruning pipelines to reduce size and latency on edge devices while preserving accuracy, publish findings, and stay at the forefront of compression techniques.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Model Compression & Quantization)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Join our AI research team to advance model compression for multimodal AI systems (LLMs and VLMs). You’ll design and implement quantization, distillation, and pruning pipelines to reduce size and latency on edge devices while preserving accuracy, publish findings, and stay at the forefront of compression techniques.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
  • •Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
  • •Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
  • •Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
  • •Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.

Pay and Benefits

Perks:Remote WorkLearning BudgetHealth Insurance

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or related field, with a strong AI R&D track record.
  • •Experience with PyTorch or equivalent frameworks.
  • •Hands-on experience with model quantization (QAT and PTQ).
  • •Knowledge distillation experience for compressing large models.
  • •Model pruning experience to reduce parameters and computation.
Experience:AIMultimodalEdge computingFintech
Education:Bachelor's
Skills:Soft skillsProblem-solving
Languages:English
Tech Stack:PyTorchQuantizationQATPTQKnowledge distillationPruningTransformersLLMsVLMsC++

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website