Principal AI Cluster Performance Validation Engineer
Austin, Seattle, Santa Clara
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Problem-solving","Debug skills","Analytical mindset","Communication","Collaboration"]Own performance validation for GPU clusters, focusing on RDMA network behavior in AI cluster environments. Conduct scalability testing, benchmarking, profiling, and bottleneck analysis to optimize throughput, latency, and congestion (RoCE/RoCE v2) and improve collective communications. Partner with hardware, software, and system architecture teams to drive performance tuning, automation, and validation, and produce clear documentation and stakeholder reporting.
Loading
Loading job details...
Preparing the role view and application actions.

