Senior Inference Engineer, GPU Kernel Optimization
NVIDIA
Santa Clara, Austin, New York, Seattle
Workplace: OnsiteFull timeUSD 184,000 - 287,500 annuallyFunction: Solutions Engineering & Sales EngineeringEducation: mastersSkills: ["Collaboration"]Drive performance to the ceiling for LLM inference by building silicon-measured GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization systems. Collaborate across compiler, kernel, hardware, and framework teams to attribute bottlenecks, generate optimization policies, and validate improvements with rigorous CUPTI/NSYS/NCU profiling for production-grade throughput and latency gains.

