Site Reliability Engineer II
Cambridge, United States
Workplace: RemoteFull timeUSD 95,000 - 171,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Collaboration","Troubleshooting","On-call readiness","Willingness to learn","Curiosity"]Build and run reliable AI inference infrastructure by automating monitoring and incident response. Improve dashboards, alerts, and SLO tracking for inference workloads using Akamai observability tooling. Develop automation and runbooks, participate in CI/CD safety and rollback processes, and support on-call rotations with blameless post-mortems. Collaborate with product engineering teams to troubleshoot issues across the stack, with opportunities to focus on GPU infrastructure and Kubernetes for serverless inference workloads.
Loading
Loading job details...
Preparing the role view and application actions.

