Site Reliability Engineer
New York
Full timeUSD 175,000 - 225,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Problem-solving","Communication","Root cause investigation","Translating technical concepts"]Join the AI Technology team to ensure the stability and scalability of an internal Agentic AI platform. Define reliability standards with SLOs, error budgets, and incident response runbooks, and own observability, incident response, and scalability. Help keep agents, gateways, LLM proxies, and RAG pipelines highly available and efficient, support users through a help channel, and improve code quality and developer practices.
Loading
Loading job details...
Preparing the role view and application actions.

