Sr. Inference Optimization Engineer (local / edge runtime)
Intel
Santa Clara, Phoenix
Workplace: HybridFull timeUSD 195,200 - 361,200 annuallyFunction: Software EngineeringExperience: 8+ yearsSkills: ["Curiosity","Problem-solving","Performance optimization"]Optimize local and edge AI inference engines so small models run fast on the hardware people actually use. You’ll profile and improve latency, throughput, and memory for llama.cpp-vulkan and vLLM, tune KV cache and continuous batching, and drive quantization strategy (GGUF/AWQ/GPTQ). Benchmark across hardware tiers, reduce CPU overhead for faster engine lifecycle, and contribute upstream patches to open-source engines.

