AI Research Engineer (Pre-training - LLM & Multi-Modal)

Tether.io
Bengaluru
Workplace: RemoteFull timeFunction: Education & TrainingEducation: bachelorsSkills: ["PyTorch","Hugging Face","LLM","Multi-modal","Distributed training","NVIDIA GPUs"]

Develop and optimize large-scale pre-training for LLMs and multi-modal models, designing architectures, tokenizers, and cross-modal alignment layers. Build robust data pipelines, curate massive datasets, and run distributed experiments to improve model intelligence, efficiency, and scalability across thousands of GPUs in a remote, globally-distributed team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Pre-training - LLM & Multi-Modal)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Develop and optimize large-scale pre-training for LLMs and multi-modal models, designing architectures, tokenizers, and cross-modal alignment layers. Build robust data pipelines, curate massive datasets, and run distributed experiments to improve model intelligence, efficiency, and scalability across thousands of GPUs in a remote, globally-distributed team.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time · Permanent
Job Function: Education & Training

Key Responsibilities

  • •Large-Scale Pre-Training: Conduct foundational pre-training for LLMs and Multi-Modal models (integrating text, vision, audio, or other modalities) on large, distributed servers equipped with multi-nodes & thousands of NVIDIA GPUs.
  • •Architecture & Alignment Innovation: Design, prototype, and scale innovative architectures, tokenizers, and cross-modal alignment layers to enhance model intelligence and multi-modal understanding.
  • •Data Strategy: Source, filter, and curate massive-scale textual and multi-modal datasets, establishing robust data pipelines for efficient pre-training.
  • •Experimental Research: Independently and collaboratively execute experiments, analyze results, and refine training methodologies for optimal performance and token efficiency.
  • •Optimization & Debugging: Investigate, debug, and eliminate bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during long training runs.

Key Requirements

  • •A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
  • •Hands-on experience contributing to large-scale LLM or Multi-Modal pre-training runs on large, distributed servers equipped with thousands of NVIDIA GPUs, ensuring scalability and impactful advancements in model performance.
  • •Familiarity and practical experience with large-scale, distributed training frameworks, libraries and tools.
  • •Deep knowledge of state-of-the-art transformer and non-transformer modifications aimed at enhancing intelligence, efficiency and scalability.
  • •Strong expertise in PyTorch and Hugging Face libraries with practical experience in model development, continual pretraining, and deployment.
Experience:FintechBlockchainDigital finance
Education:Bachelor's
Skills:PyTorchHugging FaceLLMMulti-modalDistributed trainingNVIDIA GPUs
Languages:English
Tech Stack:PyTorchHugging FaceLLMMulti-modalDistributed trainingNVIDIA GPUs

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website