Site Reliability Engineer (SRE)

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Cross-team coordination"]

Own end-to-end reliability for a fine-tuning API platform, from CI/CD and production observability to incident response and postmortems. Set Service Level Objectives for distributed training systems, design monitoring across the full training path, and improve recovery to prevent recurrence. Help harden multi-tenant isolation and resource scheduling for LoRA-based workload co-scheduling, partnering with security teams to address production vulnerabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
23 hours ago

Site Reliability Engineer (SRE)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Own end-to-end reliability for a fine-tuning API platform, from CI/CD and production observability to incident response and postmortems. Set Service Level Objectives for distributed training systems, design monitoring across the full training path, and improve recovery to prevent recurrence. Help harden multi-tenant isolation and resource scheduling for LoRA-based workload co-scheduling, partnering with security teams to address production vulnerabilities.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Define and own end-to-end reliability across CI/CD, production observability, and incident response.
  • •Develop Service Level Objectives for distributed training systems, balancing reliability and scheduling latency with development velocity.
  • •Design and implement monitoring and observability across the full training path.
  • •Lead incident response for platform issues, ensuring rapid recovery and systematic improvements.
  • •Harden multi-tenant isolation and resource scheduling for LoRA-based workload co-scheduling, and collaborate with security teams on production vulnerabilities.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •Bachelor's degree (or equivalent) in computer science, engineering, or a similar field.
  • •Experience with distributed systems, cloud infrastructure, or site reliability engineering.
  • •Proficiency writing software for reliability problems, including tooling and automation.
  • •Experience with production incident response, postmortems, and systematic reliability improvements.
  • •Strong communication skills and demonstrated coordination across engineering and research teams.
Experience:Distributed systemsCloud infrastructureSite reliability engineering
Education:Bachelor's
Skills:CommunicationCross-team coordination
Tech Stack:CI/CDObservabilityIncident responseDistributed trainingLoRAKubernetesGPU

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website