Sr. Data Center GPU Validation and Debug Engineer

AMD
Austin
Workplace: HybridFull timeUSD 152,600 - 261,600 annuallyFunction: Software EngineeringEducation: mastersSkills: ["Technical judgment","Disciplined failure isolation","Debugging","Log/trace analysis","Performance analysis"]

Validate and debug data center GPU products across the full software and hardware stack, tracing issues from applications down through firmware, kernel, compiler, libraries, and frameworks. Reproduce failures, build minimal test cases, and triage interactions across GPU, CPU, memory, networking, power, and topology. Evaluate AI, HPC, and communication workloads, establish benchmarks and regression detection, and develop automated diagnostics, stress tests, and performance suites for release qualification.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
13 hours ago

Sr. Data Center GPU Validation and Debug Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Validate and debug data center GPU products across the full software and hardware stack, tracing issues from applications down through firmware, kernel, compiler, libraries, and frameworks. Reproduce failures, build minimal test cases, and triage interactions across GPU, CPU, memory, networking, power, and topology. Evaluate AI, HPC, and communication workloads, establish benchmarks and regression detection, and develop automated diagnostics, stress tests, and performance suites for release qualification.
Location: Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Debug GPU failures including hangs, crashes, memory faults, firmware errors, and performance regressions across the full stack.
  • •Reproduce failures and reduce them to minimal, actionable test cases.
  • •Triage system-level interactions across GPU, CPU, memory, networking, power, and platform topology.
  • •Validate AI, HPC, and communication workloads across GPU platforms, firmware, and software releases.
  • •Build automated diagnostics, stress tests, performance suites, and regression infrastructure.

Pay and Benefits

Salary: USD 152,600 - 261,600 annually

Key Requirements

  • •Strong Linux and systems-engineering experience with C/C++, Python, and shell automation.
  • •Experience with data center GPUs, accelerators, or comparable high-performance systems.
  • •Understanding of GPU architecture, memory hierarchy, parallel execution, communication, and synchronization.
  • •Ability to debug across software, firmware, kernel, and hardware boundaries.
  • •Ability to analyze logs, traces, hardware telemetry, and performance counters.
Experience:Data center GPUsAcceleratorsHigh-performance systemsLinux systemsAIHPCMulti-GPUPerformance benchmarking
Education:Master's in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Skills:Technical judgmentDisciplined failure isolationDebuggingLog/trace analysisPerformance analysis
Languages:English
Tech Stack:LinuxC/C++PythonShell automationGPU architectureGPU debuggingLinux kernel driversPCIeFirmware interactionMemory managementROCm/HIPRCCLCUDAPyTorchVLLMSGLangNUMACI systemsHardware telemetryPerformance counters

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn