AI Infrastructure Engineer

Intel
Santa Clara, Austin
Workplace: HybridFull timeUSD 170,500 - 315,490 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsEducation: bachelorsSkills: ["Performance optimization","Profiling","Problem-solving","Open-source collaboration","Cross-stack debugging"]

Own end-to-end performance optimization for LLM inference on Intel next-generation GPU architectures. Profile and resolve cross-stack bottlenecks, design and integrate custom GPU kernels for attention, MoE, quantization, and operator fusions, and upstream improvements into open-source inference frameworks like vLLM, SGLang, and PyTorch. Use systematic profiling and roofline analysis to guide hardware roadmap decisions based on real GenAI workload data.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Intel
Intel
4 hours ago

AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own end-to-end performance optimization for LLM inference on Intel next-generation GPU architectures. Profile and resolve cross-stack bottlenecks, design and integrate custom GPU kernels for attention, MoE, quantization, and operator fusions, and upstream improvements into open-source inference frameworks like vLLM, SGLang, and PyTorch. Use systematic profiling and roofline analysis to guide hardware roadmap decisions based on real GenAI workload data.
Location: Santa Clara, Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own end-to-end optimization for running state-of-the-art LLMs on Intel GPUs.
  • •Profile, diagnose, and resolve cross-stack inference performance bottlenecks.
  • •Design, write, and optimize custom high-performance GPU kernels for attention, MoE, quantization, and operator fusions.
  • •Upstream architectural improvements and hardware backends into vLLM, SGLang, PyTorch, acting as a bridge between hardware teams and the open-source community.
  • •Use roofline analysis and profiling to decompose bottlenecks and partner with architecture/compiler teams to shape future GPU roadmaps.

Pay and Benefits

Salary: USD 170,500 - 315,490 annually
Equity and Bonus:Equity
Perks:Health InsuranceRetirementPaid Leave

Key Requirements

  • •Bachelors degree in Computer Science, Software Engineering, AI/ML, or related field plus 4+ years experience; or Masters plus 3+ years; or a PhD.
  • •3+ years of software engineering experience in GPU computing, AI systems, or high-performance computing (HPC).
  • •Proficiency in modern C++ and Python, comfortable modifying complex systems-level code.
  • •Hands-on kernel development/optimization experience for GPU workloads.
  • •Experience with open-source inference engines (vLLM, SGLang, PyTorch, or similar) and upstreaming improvements.
Experience:4+ yearsGPU computingAI systemsHigh-performance computing (HPC)Generative AIOpen source
Education:Bachelor's in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field
Skills:Performance optimizationProfilingProblem-solvingOpen-source collaborationCross-stack debugging
Tech Stack:C++PythonVLLMSGLangPyTorchLlama.cppTritonSYCLCUDACUTLASSGPU kernelsKV cachingMoEQuantization

Company Brief

Intel
Designs and manufactures semiconductor chips, processors, and related hardware for PCs, data centers, networking, and embedded applications, while providing software and services to accelerate computing across industries globally.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1968
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor