Principal AI Cluster Performance Validation Engineer

AMD
Austin, Seattle, Santa Clara
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Problem-solving","Debug skills","Analytical mindset","Communication","Collaboration"]

Own performance validation for GPU clusters, focusing on RDMA network behavior in AI cluster environments. Conduct scalability testing, benchmarking, profiling, and bottleneck analysis to optimize throughput, latency, and congestion (RoCE/RoCE v2) and improve collective communications. Partner with hardware, software, and system architecture teams to drive performance tuning, automation, and validation, and produce clear documentation and stakeholder reporting.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

Principal AI Cluster Performance Validation Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own performance validation for GPU clusters, focusing on RDMA network behavior in AI cluster environments. Conduct scalability testing, benchmarking, profiling, and bottleneck analysis to optimize throughput, latency, and congestion (RoCE/RoCE v2) and improve collective communications. Partner with hardware, software, and system architecture teams to drive performance tuning, automation, and validation, and produce clear documentation and stakeholder reporting.
Location: Austin, Seattle, Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Evaluate GPU cluster scalability by testing under varied workloads, cluster sizes, configurations, and networking technologies (RoCE).
  • •Develop and execute benchmarking strategies to establish baseline performance, identify bottlenecks, and drive improvements.
  • •Collaborate with hardware and software teams to optimize GPU cluster performance, including RDMA throughput, RoCE v2 congestion, latency, and collective communications.
  • •Use profiling tools and methodologies to analyze performance bottlenecks and provide actionable improvement insights.
  • •Implement performance tuning (e.g., protocol enhancements, load balancing, parallel processing optimizations), document results, and report to stakeholders.

Key Requirements

  • •Proven experience optimizing the performance of GPU clusters, including RDMA network configuration and troubleshooting.
  • •Strong understanding of GPU architectures and parallel computing concepts for clustered environments.
  • •Hands-on skills in system-level performance analysis, debugging complex HW/FW and clustered configurations.
  • •Proficiency in scripting for automation and performance analysis (e.g., Python, Bash).
  • •Linux kernel networking expertise and experience with cluster management tools/systems.
Education:Bachelor's in computer science or electrical engineering
Skills:Problem-solvingDebug skillsAnalytical mindsetCommunicationCollaboration
Tech Stack:GPURDMANICRoCERoCE v2Linux kernel networkingPythonBashHPCMachine Learning

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn