Software Engineer, Data Infrastructure

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Collaboration","Initiative","End-to-end ownership","Cross-functional teamwork","Reliability mindset"]

Join a small team building the core data infrastructure behind distributed LLM training. You’ll design and operate fault-tolerant systems for distributed compute, data orchestration, storage, and high-throughput ingestion pipelines. Build catalogs, deduplication and quality checks, search, and lifecycle traceability for petabyte-scale multimodal data. Collaborate with researchers to improve data quality, reliability, and accelerate training cycles using open-source tools like Spark, Kafka, Beam, Ray, and Delta Lake.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
23 hours ago

Software Engineer, Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join a small team building the core data infrastructure behind distributed LLM training. You’ll design and operate fault-tolerant systems for distributed compute, data orchestration, storage, and high-throughput ingestion pipelines. Build catalogs, deduplication and quality checks, search, and lifecycle traceability for petabyte-scale multimodal data. Collaborate with researchers to improve data quality, reliability, and accelerate training cycles using open-source tools like Spark, Kafka, Beam, Ray, and Delta Lake.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, build, and operate scalable, fault-tolerant infrastructure for LLM research, including distributed compute, orchestration, and storage across modalities.
  • •Develop high-throughput systems for data ingestion, processing, and transformation, including training data catalogs, deduplication, quality checks, and search.
  • •Build systems that ensure traceability, reproducibility, and robust quality control across the data lifecycle.
  • •Implement and maintain monitoring and alerting to support platform reliability and performance.
  • •Collaborate with research teams to improve data quality and accelerate training cycles.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveRelocation

Key Requirements

  • •Bachelor’s degree (or equivalent experience) in computer science, engineering, or similar.
  • •Proficiency in at least one backend language (Python or Rust).
  • •Fluency with distributed compute frameworks such as Apache Spark or Ray.
  • •Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines.
  • •Ability to operate across the stack and own end-to-end projects.
Education:Bachelor's
Skills:CollaborationInitiativeEnd-to-end ownershipCross-functional teamworkReliability mindset
Tech Stack:PythonRustApache SparkRayKafkaBeamDelta LakeDbtTerraformAirflowParquet

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website