Software Development Engineer

AMD
Bengaluru
Workplace: OnsiteFull timeINR 1,376,970 - 1,967,100 annuallyFunction: Product ManagementExperience: 7+ yearsSkills: ["Technical troubleshooting","Client support","Documentation","Collaboration","Performance tuning"]

Join AMD’s Data Center GPU organization to help operate and optimize HPC clusters and cloud-based HPC environments for AI and parallel computing. You’ll troubleshoot GPU/CPU/network/OS and Slurm bottlenecks, deploy workloads on Kubernetes and OpenStack (Ceph/Cinder/Neutron), and build automation using Python, Ansible, and Terraform. You’ll also benchmark and tune multi-node workloads (MPI/NCCL/ROCm/CUDA) while providing real-time support to high-profile client teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

Software Development Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 6 months ago

Job Summary

Join AMD’s Data Center GPU organization to help operate and optimize HPC clusters and cloud-based HPC environments for AI and parallel computing. You’ll troubleshoot GPU/CPU/network/OS and Slurm bottlenecks, deploy workloads on Kubernetes and OpenStack (Ceph/Cinder/Neutron), and build automation using Python, Ansible, and Terraform. You’ll also benchmark and tune multi-node workloads (MPI/NCCL/ROCm/CUDA) while providing real-time support to high-profile client teams.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: Product Management
Seniority: Mid level

Key Responsibilities

  • •Manage and optimize HPC clusters to ensure high availability and performance.
  • •Troubleshoot GPU, CPU, network drivers, firmware, and OS-level issues in Slurm-based environments.
  • •Deploy and manage HPC workloads in Kubernetes for AI/ML and parallel computing.
  • •Automate provisioning and operations using Ansible and Terraform, including job scheduling, monitoring, and log analysis.
  • •Benchmark and tune multi-node HPC workloads and underlying storage/networking configurations for peak performance.

Pay and Benefits

Salary: INR 1,376,970 - 1,967,100 annually

Key Requirements

  • •7+ years of experience in high-performance computing (HPC) environments with strong systems-level debugging skills.
  • •Hands-on experience with Python, Kubernetes (K8s), Slurm, OpenStack, and Ansible.
  • •Deep knowledge of drivers, troubleshooting methods, and system-level debugging for HPC stacks.
  • •Ability to manage, optimize, and troubleshoot HPC clusters and cloud-based HPC environments.
  • •Bachelor or Masters degree in Computer Engineering or Electrical/Electronics Engineering.
Experience:7+ yearsHPC
Education:
Skills:Technical troubleshootingClient supportDocumentationCollaborationPerformance tuning
Languages:En-us
Tech Stack:PythonKubernetesK8sSlurmOpenStackAnsibleTerraformOpenShiftCephCinderNeutronDockerSingularityMPINCCLROCmCUDAUCXXPMEMInfiniBand

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn