AI/ML Compiler Developer (NPU Acceleration)

AMD
Hyderabad
Workplace: OnsiteFull timeINR 3,245,970 - 4,637,100 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Performance optimization"]

Develop AI/ML compiler kernels and dataflow schedules for AMD Ryzen processors’ XDNA NPU to accelerate LLMs and Stable Diffusion networks. Design and optimize highly tuned C/C++ kernel libraries for vector processors, collaborating with research, software, and hardware teams to exploit SIMD/VLIW parallelism. Profile, validate, and test kernels across platforms using CPU models and Python/C++ operators, ensuring correctness and performance improvements through thorough profiling, documentation, and unit/integration tests.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
6 days ago

AI/ML Compiler Developer (NPU Acceleration)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 51 minutes agoStatus: Live

Job Summary

Develop AI/ML compiler kernels and dataflow schedules for AMD Ryzen processors’ XDNA NPU to accelerate LLMs and Stable Diffusion networks. Design and optimize highly tuned C/C++ kernel libraries for vector processors, collaborating with research, software, and hardware teams to exploit SIMD/VLIW parallelism. Profile, validate, and test kernels across platforms using CPU models and Python/C++ operators, ensuring correctness and performance improvements through thorough profiling, documentation, and unit/integration tests.
Location: Hyderabad
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement optimized C++ kernel libraries for NPU/GPU acceleration.
  • •Collaborate with research and software teams to integrate kernels into the existing software stack.
  • •Work with hardware engineers to understand VLIW vector core units (e.g., MAC, GeMM, and non-linear functions) and develop vectorized code leveraging SIMD and ILP.
  • •Profile and analyze kernel performance, identify bottlenecks, and optimize throughput-critical sections.
  • •Develop CPU models and C++/Python operator implementations to validate accuracy; write unit and integration tests and validate across hardware platforms.

Pay and Benefits

Salary: INR 3,245,970 - 4,637,100 annually

Key Requirements

  • •Excellent C/C++ and Python coding skills.
  • •Strong understanding of SIMD/Tensor/VLIW processor architecture to exploit parallelism.
  • •Experience developing vectorized code using SIMD and parallel computing.
  • •Familiarity with machine learning frameworks such as TensorFlow and PyTorch.
  • •Knowledge of low-level hardware concepts like cache hierarchies and memory access patterns.
Education:
Skills:Problem-solvingPerformance optimization
Languages:English
Tech Stack:CC++PythonSIMDVLIWTensorGitTensorFlowPyTorch

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn