Senior Deep Learning Software Engineer, Inference
NVIDIA
California, Texas, New York, Washington, Massachusetts
Workplace: RemoteFull timeUSD 152,000 - 287,500Function: Software EngineeringExperience: 5+ yearsSkills: ["Performance optimization","Analysis","Tuning","Software design","Cross-collaboration","Debugging","Code optimization"]Design, build, and optimize GPU-accelerated deep learning inference software for large-scale LLM and generative AI serving. Contribute to high-performance, open-source inference frameworks and NVIDIA libraries, driving performance improvements across NVIDIA accelerators from datacenter GPUs to edge SoCs. Implement and tune model serving pipelines using CUDA kernels and tools such as CUTLASS, OAI Triton, and NCCL, collaborating across cross-functional teams to ship impactful features for real deployments.

