Software Engineer, Systems Generalist

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Collaboration","Initiative","Cross-functional teamwork","Distributed systems problem-solving","Debugging"]

Build and scale the core infrastructure that powers foundation models and internal research and product teams. You’ll solve distributed systems challenges across the full technical stack, working on core infrastructure (e.g., large Kubernetes GPU clusters), data infrastructure (Spark-based pipelines and governance), and developer productivity tooling. Join a small, high-impact engineering team and collaborate directly with researchers to accelerate experiments and improve infrastructure efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
23 hours ago

Software Engineer, Systems Generalist

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Build and scale the core infrastructure that powers foundation models and internal research and product teams. You’ll solve distributed systems challenges across the full technical stack, working on core infrastructure (e.g., large Kubernetes GPU clusters), data infrastructure (Spark-based pipelines and governance), and developer productivity tooling. Join a small, high-impact engineering team and collaborate directly with researchers to accelerate experiments and improve infrastructure efficiency.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Architect and scale core infrastructure systems that support training, research, and serving of AI foundation models.
  • •Design and optimize data pipelines and data infrastructure (including governance best practices) for research and products.
  • •Build and maintain large-scale cluster workloads reliably and safely, including Kubernetes-based GPU workloads.
  • •Develop tooling and systems that improve developer productivity, configuration, and optimized environments.
  • •Collaborate with researchers to accelerate experiments, improve infrastructure efficiency, and enable insights across models, products, and data assets.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveParental LeaveRelocation

Key Requirements

  • •Bachelor’s degree (or equivalent) in computer science, engineering, or a related field.
  • •Proficiency in at least one backend language (Python or Rust).
  • •Experience operating large-scale clusters and container orchestration systems (e.g., Kubernetes or Slurm).
  • •Comfort operating across the stack and owning projects end-to-end.
  • •Thrive in a highly collaborative, cross-functional environment; show a bias for action and initiative.
Experience:AIFoundation modelsInfrastructureDistributed systems
Education:Bachelor's
Skills:CollaborationInitiativeCross-functional teamworkDistributed systems problem-solvingDebugging
Tech Stack:PythonRustKubernetesSlurmSparkContainersCIGPU/MLGPU workloadsTinker

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website