Fellow Software Engineer — AI Performance & Reliability

AMD
San Jose, Bellevue
Workplace: HybridFull timeFunction: Software EngineeringEducation: phdSkills: ["Written communication","Verbal communication","Debugging","Analytical skills","Customer-focused mindset"]

Join the AI Infrastructure team to improve the performance, efficiency, and reliability of AI workloads across both model training and inference. Work on large language models, diffusion models, and recommendation systems by identifying bottlenecks across the stack, developing reusable solutions, and collaborating with customers and engineering teams. Build performance tooling, benchmarks, and observability systems, and help resolve complex production issues at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 month ago

Fellow Software Engineer — AI Performance & Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Join the AI Infrastructure team to improve the performance, efficiency, and reliability of AI workloads across both model training and inference. Work on large language models, diffusion models, and recommendation systems by identifying bottlenecks across the stack, developing reusable solutions, and collaborating with customers and engineering teams. Build performance tooling, benchmarks, and observability systems, and help resolve complex production issues at scale.
Location: San Jose, Bellevue
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Improve the performance, efficiency, and reliability of AI workloads across model training and inference.
  • •Identify and optimize bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  • •Optimize workloads for large language models, diffusion models, recommendation systems, and other ML architectures.
  • •Develop performance tooling, benchmarks, automation, and observability systems.
  • •Collaborate with customers and internal teams to reproduce issues, resolve production problems, and translate feedback into product and infrastructure improvements.

Key Requirements

  • •Build production-quality software systems with strong software engineering skills.
  • •Work with AI infrastructure for model training and/or inference.
  • •Profile and optimize machine learning models and AI workloads to improve throughput, latency, memory efficiency, scalability, and reliability.
  • •Apply strong computer architecture foundations, including processors, memory hierarchies, and parallelism.
  • •Use systems-oriented programming languages such as Python or C++, and work with ML frameworks like PyTorch, TensorFlow, or JAX.
Experience:Machine learningAI infrastructureLarge language modelsDiffusion modelsHigh-performance computing
Education:PhD / Doctorate in artificial intelligence, machine learning, computer science, or a related field
Skills:Written communicationVerbal communicationDebuggingAnalytical skillsCustomer-focused mindset
Languages:En-us
Tech Stack:PythonC++PyTorchTensorFlowJAXROCmHIPCUDATritonXLAMLIRNCCLGPUAcceleratorDistributed computingBenchmarksAutomationObservabilityProfiling

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn