AI Infrastructure Engineer

Thinking Machines Lab
San Francisco, New York
Workplace: OnsiteFull timeUSD 350,000 - 475,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Ownership","Collaboration","Judgment"]

Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, partnering with research teams during active runs to debug failures across accelerators, networking, storage, schedulers, and training frameworks. Build monitoring, alerting, automated recovery, and improvements to checkpointing, fault tolerance, and job scheduling. Develop internal tooling to reduce toil and improve cluster utilization, and participate in on-call with postmortems that drive permanent fixes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 minutes agoStatus: Live

Job Summary

Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, partnering with research teams during active runs to debug failures across accelerators, networking, storage, schedulers, and training frameworks. Build monitoring, alerting, automated recovery, and improvements to checkpointing, fault tolerance, and job scheduling. Develop internal tooling to reduce toil and improve cluster utilization, and participate in on-call with postmortems that drive permanent fixes.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability, performance, and uptime of large-scale post-training and RL training jobs from launch through completion.
  • •Partner directly with research teams during active model runs to unblock training and speed iteration.
  • •Debug failures across accelerators, networking, storage, schedulers, and training frameworks to root cause.
  • •Build monitoring, alerting, and automated recovery so runs self-heal or fail fast.
  • •Improve checkpointing, fault tolerance, and job scheduling, and build internal tools to reduce toil and improve cluster utilization.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production.
  • •Track record debugging complex failures across distributed systems (networking, hardware, kernel, scheduler).
  • •Strong software engineering skills in Python and/or Go/C++, with judgment to decide when to script vs. build systems.
  • •Solid grounding in Linux systems internals and networking fundamentals.
  • •Comfortable owning production systems, including participating in on-call rotations.
Experience:Distributed systemsMachine learningAI trainingResearch environments
Skills:Problem-solvingOwnershipCollaborationJudgment
Tech Stack:PythonGoC++LinuxPyTorchRaySlurmKubernetesInfiniBandRDMANCCLGPUTPU

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website