Software Engineer, Supercomputing

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Collaboration","Initiative","Cross-functional communication","End-to-end ownership"]

Design, build, and operate the GPU supercomputing environment that enables large-scale AI training and inference. Run and automate GPU clusters (provisioning, imaging, capacity planning), build software to abstract cluster management, and extend orchestration for efficient scheduling and multi-tenancy. Monitor performance and reliability, improve error recovery, and support researchers with scalable training and performance trade-offs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
20 hours ago

Software Engineer, Supercomputing

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Design, build, and operate the GPU supercomputing environment that enables large-scale AI training and inference. Run and automate GPU clusters (provisioning, imaging, capacity planning), build software to abstract cluster management, and extend orchestration for efficient scheduling and multi-tenancy. Monitor performance and reliability, improve error recovery, and support researchers with scalable training and performance trade-offs.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Operate and automate large GPU clusters, including provisioning, imaging, and capacity planning.
  • •Develop software to abstract cluster management and provide a unified interface for training and inference.
  • •Extend scheduling/orchestration (Kubernetes, Slurm, or similar) for topology-aware placement, preemption, quotas, and fair-share multi-tenancy.
  • •Monitor and improve operational metrics for speed, reliability, and error recovery.
  • •Build reliable storage and artifact paths for datasets, checkpoints, and logs with retention and lineage, and partner with researchers to unblock scale runs.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • •Proficiency in at least one backend language (Python or Rust).
  • •Experience operating large-scale clusters and container orchestration systems (Kubernetes or Slurm).
  • •Comfort operating across the stack and owning projects end-to-end.
  • •Thrive in a collaborative environment and take initiative across teams to ensure delivery.
Education:Bachelor's
Skills:CollaborationInitiativeCross-functional communicationEnd-to-end ownership
Tech Stack:PythonRustKubernetesSlurmLinuxNetworkingInfrastructure-as-codeCUDANCCLPyTorchTensorFlowJAXDeep learningGPUDistributed trainingInference

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website