Site Reliability Engineer, Production

Thinking Machines Lab
San Francisco, New York
Workplace: OnsiteFull timeUSD 350,000 - 475,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Coordination"]

Own end-to-end reliability for the Tinker platform, from CI/CD through production observability and incident response. Define SLOs for distributed training workloads, build monitoring and recovery systems, and lead systematic improvements to prevent recurrence. Harden multi-tenant isolation and resource scheduling for efficient LoRA co-scheduling, while partnering with security teams to address production vulnerabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 month ago

Site Reliability Engineer, Production

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Own end-to-end reliability for the Tinker platform, from CI/CD through production observability and incident response. Define SLOs for distributed training workloads, build monitoring and recovery systems, and lead systematic improvements to prevent recurrence. Harden multi-tenant isolation and resource scheduling for efficient LoRA co-scheduling, while partnering with security teams to address production vulnerabilities.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define and own end-to-end reliability across CI/CD flows, production observability, and incident response.
  • •Develop SLOs for distributed training systems, balancing job completion reliability and scheduling latency.
  • •Design and implement monitoring and observability across the full training path.
  • •Drive incident response and incident reviews for Tinker platform issues, ensuring rapid recovery and prevention of recurrence.
  • •Harden multi-tenant isolation and resource scheduling so LoRA-based workloads can be co-scheduled without compromising reliability or data separation.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • •Experience in distributed systems, cloud infrastructure, or site reliability engineering.
  • •Proficiency writing software to solve reliability problems, including building tooling and automation.
  • •Experience with production incident response, postmortems, and systematic reliability improvement.
  • •Strong communication skills and a track record of coordinating across engineering and research teams.
Experience:Distributed systemsCloud infrastructureSite reliability engineeringProduction incident responseDistributed training
Education:Bachelor's
Skills:CommunicationCoordination
Tech Stack:CI/CDProduction observabilityIncident responseKubernetesLoRAService level objectives

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website