Senior Software Engineer - AI Inference Performance
Santa Clara
Workplace: OnsiteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 6+ yearsEducation: mastersSkills: ["Hands-on","Analytical","Collaborative","Autonomous","Creative"]Advance LLM/VLM inference performance on NVIDIA GPU-accelerated systems by turning profiler and performance models into production-ready code. Lead end-to-end analysis of prefill/decode workloads, optimize latency and throughput metrics, and benchmark against reproducible regression gates. Profile with NVIDIA Nsight tools and PyTorch Profiler, tune serving and KV-cache strategies, and build performance-critical CUDA/CUTLASS/Triton kernels. Collaborate across teams and improve open-source inference engines.
Loading
Loading job details...
Preparing the role view and application actions.

