Member of Technical Staff - GPU Performance Engineer
San Francisco, Boston
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["CUDA","C/C++","Nsight","PyTorch","GPU kernels","CUDA kernels"]Seeking a highly autonomous GPU performance engineer to design, implement, and optimize custom CUDA kernels for cutting-edge AI models. You’ll profile at the hardware level, integrate kernels into PyTorch pipelines, and collaborate with researchers to ship speedups in training, post-training, and inference. The role emphasizes memory hierarchies, tensor cores, and end-to-end performance benchmarks in a small, high-ownership team.

