Site Reliability Engineer, Production
San Francisco, New York
Workplace: OnsiteFull timeUSD 350,000 - 475,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Coordination"]Own end-to-end reliability for the Tinker platform, from CI/CD through production observability and incident response. Define SLOs for distributed training workloads, build monitoring and recovery systems, and lead systematic improvements to prevent recurrence. Harden multi-tenant isolation and resource scheduling for efficient LoRA co-scheduling, while partnering with security teams to address production vulnerabilities.
Loading
Loading job details...
Preparing the role view and application actions.

