Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Education & TrainingEducation: bachelorsSkills: ["Collaboration","Debugging","Initiative","Performance optimization","Cross-functional teamwork"]

Design and build scalable training infrastructure for large models, owning the training stack end to end. You’ll implement distributed training systems that scale across thousands of GPUs, optimize for throughput and reliability, and create reusable frameworks that improve reproducibility and security. Partner with researchers and engineers to remove system bottlenecks and publish internal documentation and technical reports that advance scalable AI infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

Research Engineer, Infrastructure, Training Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design and build scalable training infrastructure for large models, owning the training stack end to end. You’ll implement distributed training systems that scale across thousands of GPUs, optimize for throughput and reliability, and create reusable frameworks that improve reproducibility and security. Partner with researchers and engineers to remove system bottlenecks and publish internal documentation and technical reports that advance scalable AI infrastructure.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.
  • •Develop high-performance optimizations to maximize throughput and efficiency.
  • •Build reusable frameworks and libraries to improve training reproducibility, reliability, and scalability.
  • •Set standards for reliability, maintainability, and security for systems under rapid iteration.
  • •Collaborate with researchers and engineers and publish learnings via documentation, open-source libraries, or technical reports.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •Bachelor’s degree (or equivalent experience) in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • •Strong engineering skills to write performant, maintainable code and debug complex codebases.
  • •Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • •Ability to thrive in a highly collaborative, cross-functional environment with many partners and subject-matter experts.
  • •Mindset to take initiative across different stacks and teams to ensure work ships.
Education:Bachelor's
Skills:CollaborationDebuggingInitiativePerformance optimizationCross-functional teamwork
Tech Stack:PyTorchJAXXLAMegatron-LMDeepSpeed

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website