AI Research Engineer (Pre-training - LLM & Multi-Modal)

Tether.io
Spain
Workplace: RemoteFull timeFunction: Education & TrainingEducation: high_schoolSkills: ["PyTorch","Hugging Face","Distributed training","Transformers","NVIDIA GPUs","Python"]

Join a global, remote AI model team to drive foundational pre-training for LLMs and multi-modal systems. You will design scalable architectures, curate large-scale datasets, and optimize training pipelines across distributed GPU clusters, pushing state-of-the-art in cross-modal AI performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
3 months ago

AI Research Engineer (Pre-training - LLM & Multi-Modal)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Join a global, remote AI model team to drive foundational pre-training for LLMs and multi-modal systems. You will design scalable architectures, curate large-scale datasets, and optimize training pipelines across distributed GPU clusters, pushing state-of-the-art in cross-modal AI performance.
Location: Spain
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Large-Scale Pre-Training: Conduct foundational pre-training for LLMs and multi-modal models on large distributed servers with thousands of NVIDIA GPUs.
  • •Architecture & Alignment Innovation: Design, prototype, and scale innovative architectures, tokenizers, and cross-modal alignment layers.
  • •Data Strategy: Source, filter, and curate massive-scale textual and multi-modal datasets, establishing robust data pipelines for efficient pre-training.
  • •Experimental Research: Independently and collaboratively execute experiments, analyze results, and refine training methodologies for optimal performance and token efficiency.
  • •Optimization & Debugging: Investigate, debug, and eliminate bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during long training runs.
  • •System Scalability: Contribute to the advancement of distributed training systems to ensure seamless scalability and hardware efficiency on target platforms.

Key Requirements

  • •Degree in Computer Science or related field; PhD in NLP/ML ideal with a strong AI R&D track record
  • •Hands-on experience with large-scale LLM or multi-modal pre-training on distributed servers with thousands of GPUs
  • •Familiarity with large-scale distributed training frameworks and tools
  • •Deep knowledge of transformer and non-transformer model modifications for efficiency and scalability
  • •Strong expertise in PyTorch and Hugging Face libraries; practical model development, pretraining, and deployment
Experience:AILLMMultimodalNLPDistributed training
Education:High School
Skills:PyTorchHugging FaceDistributed trainingTransformersNVIDIA GPUsPython
Languages:English
Tech Stack:PyTorchHugging FaceNVIDIA GPUsDistributed trainingTransformersPython

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website