Principal AI Solutions Architect
Santa Clara
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Problem-solving","Collaboration"]Lead design and deployment of production-grade Kubernetes-based AI infrastructure on AMD GPUs, enabling large-scale LLM training and inference. Architect multi-node, multi-datacenter GPU clusters with topology-aware scheduling, SLURM integration, and Kubernetes controllers (Kubeflow, MPI Operator, Volcano, Kueue). Translate advanced inference frameworks (vLLM, SGLang) into customer-ready solutions, accelerating time-to-production and optimizing performance.
Loading
Loading job details...
Preparing the role view and application actions.

