AI Systems Engineer - HPC
San Jose
Workplace: OnsiteFull timeUSD 151,900 - 260,400 annuallyFunction: IT Operations (Systems/Network Admin)Skills: ["Organizational skills","Problem-solving","Troubleshooting","Communication"]Design, develop, and administer GPU-focused HPC infrastructure for AMD’s AI compute platforms. Build and maintain GPU-based clusters, manage deployments, and administer distributed ML services including LLMs and AI inferencing. Automate provisioning and cluster management, collaborate with cross-functional teams on AI infrastructure needs, and monitor performance to ensure best-practice reliability and efficiency. Use AI/ML to improve internal delivery tools and processes.
Loading
Loading job details...
Preparing the role view and application actions.

