Site Reliability Engineer (SRE)
Thinking Machines Lab
San Francisco
Workplace: OnsiteFull time350,000 - 475,000Function: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Problem-solving","Collaboration","Analytical thinking"]Site Reliability Engineer to drive end-to-end reliability for the Tinker platform, collaborating with platform engineers and research teams to build robust monitoring, incident response, and multi-tenant scheduling. You’ll define SLOs for distributed training systems, design observability across the full training path, and improve reliability while maintaining data separation.

