Senior Site Reliability Engineer
Indiana, India
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Collaboration","Automation","Incident response","Cross-functional coordination","Ownership"]Join the AI Hardware SRE team to oversee, scale, and optimize next-generation dedicated AI hardware infrastructure. Define KPIs, drive proactive monitoring, automation, and rapid incident resolution across high-density hardware and regional data centers. Build Python-based infrastructure-as-code and operational tooling, implement observability with Prometheus/Grafana and OpenTelemetry/Loki, and lead telemetry/telemetry baselines for reliable service rollouts while participating in 24x7x365 on-call.
Loading
Loading job details...
Preparing the role view and application actions.

