Inference Optimization Engineer (local / edge runtime)
Intel
Santa Clara, Phoenix
Workplace: HybridFull timeUSD 170,500 - 315,490 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Problem-solving","Communication","Collaboration"]Develop and optimize local/edge AI inference engines (llama.cpp, vLLM) for Intel edge devices, focusing on latency, throughput, and memory on edge GPUs/CPUs. You’ll tune KV cache, batching, quantization, and scheduling, cut CPU overhead, bench across hardware tiers, and upstream fixes to open-source engines—enabling high-performance, private AI at the edge.

