Staff Software Engineer - GenAI inference

Databricks
San Francisco
Workplace: OnsiteFull timeUSD 190,900 - 232,800 annuallyFunction: Software EngineeringExperience: 6+ yearsEducation: mastersSkills: ["Communication","Leadership","Ownership","Collaboration","Proactive"]

Lead architecture, development, and optimization of the GenAI inference engine powering the Databricks Foundation Model API. Drive high throughput, low latency, and scalable inference across GPUs and accelerators, collaborating with researchers to bring new models and features, while building instrumentation and tooling for profiling and reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
10 months ago

Staff Software Engineer - GenAI inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead architecture, development, and optimization of the GenAI inference engine powering the Databricks Foundation Model API. Drive high throughput, low latency, and scalable inference across GPUs and accelerators, collaborating with researchers to bring new models and features, while building instrumentation and tooling for profiling and reliability.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
  • •Partner with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
  • •Lead end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs and accelerators
  • •Define and guide standards to build instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
  • •Architect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads

Pay and Benefits

Salary: USD 190,900 - 232,800 annually

Key Requirements

  • •BS/MS/PhD in Computer Science, or a related field
  • •Strong software engineering background (6+ years or equivalent) in performance-critical systems
  • •Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
  • •Strong background in distributed systems design, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
  • •Experience building instrumentation, tracing, and profiling tools for ML models
Experience:6+ yearsAIMLGPU
Education:Master's
Skills:CommunicationLeadershipOwnershipCollaborationProactive
Languages:English
Tech Stack:CUDACuBLASCuDNNNCCLGPUsDistributed systemsRPCModel serving

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn