Member of Technical Staff - GPU Performance Engineer

Liquid AI
San Francisco, Boston, United States, Cambridge
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["CUDA","C/C++","Nsight","PyTorch","GPU kernels","CUDA kernels"]

Seeking a highly autonomous GPU performance engineer to design, implement, and optimize custom CUDA kernels for cutting-edge AI models. You’ll profile at the hardware level, integrate kernels into PyTorch pipelines, and collaborate with researchers to ship speedups in training, post-training, and inference. The role emphasizes memory hierarchies, tensor cores, and end-to-end performance benchmarks in a small, high-ownership team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Liquid AI
Liquid AI
1 year ago

Member of Technical Staff - GPU Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Seeking a highly autonomous GPU performance engineer to design, implement, and optimize custom CUDA kernels for cutting-edge AI models. You’ll profile at the hardware level, integrate kernels into PyTorch pipelines, and collaborate with researchers to ship speedups in training, post-training, and inference. The role emphasizes memory hierarchies, tensor cores, and end-to-end performance benchmarks in a small, high-ownership team.
Location: San Francisco, Boston, United States, Cambridge
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Write high-performance GPU kernels for our novel model architectures
  • •Integrate kernels into PyTorch pipelines (custom ops, extensions, dispatch, benchmarking)
  • •Profile and optimize training and inference workflows to eliminate bottlenecks
  • •Build correctness tests and numerics checks
  • •Build/maintain performance benchmarks and guardrails to prevent regressions

Pay and Benefits

Perks:Health Insurance401kPaid LeaveVision

Key Requirements

  • •Authored custom CUDA kernels (not only calling cuDNN/cuBLAS)
  • •Strong understanding of GPU architecture and performance: memory hierarchy, warps, shared memory/register pressure, bandwidth vs compute limits
  • •Proficiency with low-level profiling (Nsight Systems/Compute) and performance methodology
  • •Strong C/C++ skills
  • •Experience with PyTorch custom ops or similar kernel integration (Nice-to-have)
Experience:GPU computingAIDeep learning
Skills:CUDAC/C++NsightPyTorchGPU kernelsCUDA kernels
Tech Stack:CUDANsight SystemsNsight ComputePyTorchC++C

Company Brief

Liquid AI
Builds AI infrastructure and tooling to enable real-time, distributed machine learning and orchestration across edge and cloud environments, simplifying deployment and management of intelligent applications.
Industry: AI & Machine Learning
Website