HPC Systems Engineer - AI Workloads
San Jose
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Education: bachelorsSkills: ["Organizational skills","Problem-solving","Troubleshooting","Verbal communication","Written communication"]Design, develop, and administer HPC infrastructure for AI workloads, focusing on GPU clusters and AI workload schedulers. Build and optimize GPU-based cluster performance, administer distributed ML/LLM and AI inferencing platforms, and automate system provisioning and cluster management. Partner cross-functionally to meet AI infrastructure needs, while monitoring and improving performance using best practices and tools such as Prometheus/Grafana and automation frameworks.
Loading
Loading job details...
Preparing the role view and application actions.

