AI Infrastructure Engineer
42dot
South Korea
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Problem-solving"]Operate and maintain a large-scale GPU cluster spanning multiple data centers, using Kubernetes and Slurm to orchestrate thousands of GPUs. Monitor and troubleshoot GPU hardware and software issues to ensure high availability and fast recovery. Build automation tools with Python or Shell, manage GPU quota for ML teams, and support the architecture and performance tuning of distributed training environments for large-scale autonomous driving models.

