Fellow, AI Workload Optimization

AMD
Bellevue
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 15+ yearsEducation: mastersSkills: ["Communication","Leadership","Problem-solving","Mentorship"]

visionary technical leader to define and drive end-to-end software optimization strategy for AI workloads, aligning architecture, customer engagement, and software engineering to maximize performance of AMD’s ROCm, compilers, and AI frameworks across large-scale models in data-center environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
3 months ago

Fellow, AI Workload Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

visionary technical leader to define and drive end-to-end software optimization strategy for AI workloads, aligning architecture, customer engagement, and software engineering to maximize performance of AMD’s ROCm, compilers, and AI frameworks across large-scale models in data-center environments.
Location: Bellevue
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Define and communicate the technical vision and roadmap for workload optimization across the AI software stack to attract top-tier AI customers.
  • •Lead profiling, analysis, and tuning of large-scale models (LLMs, Diffusion, Multimodal, MoE) to ensure out-of-the-box performance on AMD hardware.
  • •Partner with customers and hyperscalers to understand workload requirements and deliver architectural wins and software optimizations.
  • •Collaborate across hardware architecture, compiler, and framework teams to influence future silicon features based on AI workload trends.
  • •Drive ecosystem innovation by developing tools and frameworks for performance estimation, modeling, and automated reporting.

Key Requirements

  • •15+ years of software development experience with at least 5 years in a high-level technical leadership role (Fellow or equivalent).
  • •Deep expertise in AI frameworks (PyTorch, JAX, vLLM, SGLang) and the ROCm software stack.
  • •Proven history of optimizing distributed inference and training at scale across multi-node/multi-GPU environments.
  • •Mastery of performance profiling tools (e.g., TorchProfiler, ROCm Profiler, Nsight) and hardware-level performance modeling.
  • •Strong understanding of modern model architectures (Transformer, Attention, KV Cache) and optimization techniques like quantization, speculative decoding, and FlashAttention.
Experience:15+ yearsAIAccelerated computingData centersMachine learningHigh performance computing
Education:Master's
Skills:CommunicationLeadershipProblem-solvingMentorship
Languages:English
Tech Stack:PyTorchJAXVLLMSGLangROCmTorchProfilerNsightTransformerFlashAttentionProfiling

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn