AI Infrastructure Engineer

42dot
South Korea
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Problem-solving"]

Operate and maintain a large-scale GPU cluster spanning multiple data centers, using Kubernetes and Slurm to orchestrate thousands of GPUs. Monitor and troubleshoot GPU hardware and software issues to ensure high availability and fast recovery. Build automation tools with Python or Shell, manage GPU quota for ML teams, and support the architecture and performance tuning of distributed training environments for large-scale autonomous driving models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
42dot
42dot
1 month ago

AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Operate and maintain a large-scale GPU cluster spanning multiple data centers, using Kubernetes and Slurm to orchestrate thousands of GPUs. Monitor and troubleshoot GPU hardware and software issues to ensure high availability and fast recovery. Build automation tools with Python or Shell, manage GPU quota for ML teams, and support the architecture and performance tuning of distributed training environments for large-scale autonomous driving models.
Location: South Korea
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and maintain a large-scale GPU cluster across multiple data centers using Kubernetes and Slurm.
  • •Monitor and diagnose failures across GPU hardware and software stacks to ensure high availability and rapid recovery.
  • •Develop automation tools and scripts in Python or Shell to streamline repetitive infrastructure management.
  • •Manage GPU resource quotas and provide technical support to ML researchers for optimal computing utilization.
  • •Participate in architectural design and performance tuning for distributed training environments for large-scale autonomous driving models.

Key Requirements

  • •Strong understanding of Linux, including kernel operations, process management, and system security.
  • •Practical experience with containerization and orchestration using Docker and Kubernetes.
  • •Understanding of networking fundamentals (TCP/IP, HTTP(S)) with basic troubleshooting ability.
  • •Ability to write maintainable automation/system administration scripts in Python or Shell.
  • •Logical problem-solving skills to identify and resolve root causes in complex, large-scale systems.
Experience:AI infrastructureGPU clustersDistributed trainingAutonomous driving modelsDeep learning
Skills:CommunicationProblem-solving
Tech Stack:LinuxKubernetesSlurmGPUDockerPythonShellTCP/IPHTTP(S)PrometheusGrafanaDatadogAWSGCPCUDANCCLPyTorchTensorFlowTerraform

Company Brief

42dot
Develops autonomous driving and mobility software platforms, including AI-based perception, mapping, routing, and connected-vehicle technologies. It works on next-generation transportation systems and self-driving vehicle capabilities for automotive applications.
Industry: Autonomous Vehicles
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Seoul, South Korea
Founded: 2019
WebsiteLinkedIn