Staff Software Engineer - GenAI Performance and Kernel

Databricks
San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: mastersSkills: ["Communication","Leadership","Mentoring","Collaboration","Problem-solving"]

Staff Software Engineer for GenAI Performance and Kernel leads design, implementation, and optimization of high-performance GPU kernels powering the GenAI inference stack. You will drive kernel-level performance, balance hardware efficiency with generality, mentor engineers, and collaborate with ML researchers, systems engineers, and product teams to push inference performance at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
10 months ago

Staff Software Engineer - GenAI Performance and Kernel

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Staff Software Engineer for GenAI Performance and Kernel leads design, implementation, and optimization of high-performance GPU kernels powering the GenAI inference stack. You will drive kernel-level performance, balance hardware efficiency with generality, mentor engineers, and collaborate with ML researchers, systems engineers, and product teams to push inference performance at scale.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Lead the design, implementation, benchmarking, and maintenance of core compute kernels (e.g. attention, MLP, softmax, layernorm, memory management) optimized for various hardware backends (GPU, accelerators)
  • •Drive the performance roadmap for kernel-level improvements: vectorization, tensorization, tiling, fusion, mixed precision, sparsity, quantization, memory reuse, scheduling, auto-tuning, etc.
  • •Integrate kernel optimizations with higher-level ML systems
  • •Build and maintain profiling, instrumentation, and verification tooling to detect correctness, performance regressions, numerical issues, and hardware utilization gaps
  • •Lead performance investigations and root-cause analysis on inference bottlenecks, e.g. memory bandwidth, cache contention, kernel launch overhead, tensor fragmentation

Key Requirements

  • •BS/MS/PhD in Computer Science, or a related field
  • •Deep hands-on experience writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar sort) for ML workloads
  • •Strong knowledge of GPU/accelerator architecture: warp structure, memory hierarchy (global, shared, register, L1/L2 caches), tensor cores, scheduling, SM occupancy, etc.
  • •Experience with advanced optimization techniques: tiling, blocking, software pipelining, vectorization, fusion, loop transformations, auto-tuning
  • •Familiarity with ML-specific kernel libraries (cuBLAS, cuDNN, CUTLASS, oneDNN, etc.) or open kernels
Experience:AIGPU computingCloud computingMachine learningSystems software
Education:Master's
Skills:CommunicationLeadershipMentoringCollaborationProblem-solving
Languages:English
Tech Stack:CUDATritonOpenCLLLVM IRAssemblyCuBLASCuDNNCUTLASSOneDNNNsightNVProfPerfVtune

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn