Software Engineer, ML Infra

Thinking Machines Lab
San Francisco, New York
Workplace: OnsiteFull timeUSD 350,000 - 475,000 annuallyFunction: Software EngineeringSkills: ["Judgment","Incident response","Mentoring","Cross-stack troubleshooting","Researcher communication"]

Build and operate ML infrastructure as the day-to-day interface between research and systems. Debug issues end-to-end across kernel, NCCL, scheduler, applications, and telemetry, and provide hands-on support during major “hero run” incidents. Act as the front door for researchers when ownership isn’t clear, lead postmortems, and develop tooling to prevent recurrence while mentoring engineers across cross-stack breadth.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

Software Engineer, ML Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build and operate ML infrastructure as the day-to-day interface between research and systems. Debug issues end-to-end across kernel, NCCL, scheduler, applications, and telemetry, and provide hands-on support during major “hero run” incidents. Act as the front door for researchers when ownership isn’t clear, lead postmortems, and develop tooling to prevent recurrence while mentoring engineers across cross-stack breadth.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Debug issues across the full stack (kernel, NCCL, scheduler, application, telemetry) to find root causes.
  • •Provide embedded hands-on support during hero runs and major incidents until resolved.
  • •Act as the front door for researchers when something breaks and ownership isn’t obvious.
  • •Lead postmortems and build tooling to prevent the next incident.
  • •Mentor other engineers to develop cross-stack breadth and effective escalation judgment.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •Hands-on competence in 4+ areas including Linux kernel, networking, GPUs/CUDA, distributed systems runtimes, storage, and compilers/language runtimes, plus observability internals.
  • •Comfort operating without a clearly defined owner and strong judgment on when to dig in vs. when to escalate.
  • •Track record of being the person other engineers escalate to across 2+ companies.
  • •Shipped meaningful contributions in 3+ distinct technical stacks.
  • •Experience operating at frontier training/inference cluster scale and willingness to own the hardest, least-defined problems.
Experience:AI/ML infrastructureDistributed systemsFrontier trainingInference clusters
Skills:JudgmentIncident responseMentoringCross-stack troubleshootingResearcher communication
Tech Stack:Linux kernelNetworkingGPUsCUDANCCLDistributed systems runtimesStorageCompilersLanguage runtimesObservabilityTelemetryScheduler

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website