AI Research Engineer (Multi-Modal & Vision)

Tether.io
Bengaluru
Workplace: RemoteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Research","Engineering","Distributed training","Multimodal","Vision-language","Dataset curation"]

Join an AI model team focused on training and optimizing vision-language models for production. You’ll handle end-to-end research and engineering across training, evaluation, and deployment, build multimodal datasets, improve model efficiency for resource-constrained environments, create evaluation frameworks, scale distributed GPU training, and contribute to open-source tooling while publishing in top venues.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
2 months ago

AI Research Engineer (Multi-Modal & Vision)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join an AI model team focused on training and optimizing vision-language models for production. You’ll handle end-to-end research and engineering across training, evaluation, and deployment, build multimodal datasets, improve model efficiency for resource-constrained environments, create evaluation frameworks, scale distributed GPU training, and contribute to open-source tooling while publishing in top venues.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Conduct end-to-end research and engineering on vision-language models, covering training, evaluation, and optimization across the full model development lifecycle.
  • •Design and implement post-training pipelines including supervised fine-tuning, knowledge distillation, and reinforcement learning from human feedback.
  • •Develop and maintain high-quality multimodal datasets, including data curation, filtering, and balancing for domain-specific tasks.
  • •Drive model efficiency and deployability, adapting models for resource-constrained environments using compression and optimization techniques.
  • •Design and implement evaluation frameworks and benchmarks to measure model performance, robustness, and real-world task success.

Key Requirements

  • •Degree in Computer Science, Machine Learning, or a related field; MS/PhD preferred.
  • •Strong experience with multimodal post-training workflows including supervised fine-tuning, knowledge distillation, and reinforcement learning from feedback.
  • •Hands-on experience with parameter-efficient fine-tuning and distributed training frameworks.
  • •Demonstrated ability to build and improve vision-language models with measurable results on standard benchmarks or real-world tasks.
  • •Experience adapting models for resource-constrained environments.
Experience:AiMultimodalVision
Education:Bachelor's
Skills:ResearchEngineeringDistributed trainingMultimodalVision-languageDataset curation
Languages:English
Tech Stack:PythonPyTorchDistributed trainingGPUDistillationFine-tuningQuantizationVision-language models

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website