Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara
Workplace: HybridFull timeUSD 124,000 - 195,500 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Experimental methodology","Communication","Optimization","Benchmarking","Collaboration"]Drive performance improvements for large-scale AI inference by designing and validating autonomous, evidence-backed optimization workflows. Explore configuration options to increase throughput-per-GPU and interactivity while handling batching, KV cache, quantization, and speculative decoding. Benchmark and profile across serving architectures and platforms using Nsight tools, CUPTI, and profiling analysis, then upstream serving patches, optimized kernels, and deployment recipes.
Loading
Loading job details...
Preparing the role view and application actions.

