Site Reliability Engineer (SRE)
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Cross-team coordination"]Own end-to-end reliability for a fine-tuning API platform, from CI/CD and production observability to incident response and postmortems. Set Service Level Objectives for distributed training systems, design monitoring across the full training path, and improve recovery to prevent recurrence. Help harden multi-tenant isolation and resource scheduling for LoRA-based workload co-scheduling, partnering with security teams to address production vulnerabilities.
Loading
Loading job details...
Preparing the role view and application actions.

