Staff Software Engineer, Model Serving

Databricks
San Francisco, Mountain View
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 10+ yearsSkills: ["Communication","Leadership","Problem-solving","Mentoring","Team collaboration"]

Lead the design and delivery of high-throughput, low-latency model serving infrastructure for CPU and GPU workloads. You’ll shape architecture, collaborate across platform, product, infrastructure, and research teams, and mentor engineers while driving performance, scalability, and cost-efficiency for the Databricks AI platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
10 months ago

Staff Software Engineer, Model Serving

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead the design and delivery of high-throughput, low-latency model serving infrastructure for CPU and GPU workloads. You’ll shape architecture, collaborate across platform, product, infrastructure, and research teams, and mentor engineers while driving performance, scalability, and cost-efficiency for the Databricks AI platform.
Location: San Francisco, Mountain View
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and implement core systems and APIs that power DatabricksModel Serving, ensuring scalability, reliability, and operational excellence.
  • •Partner with product and engineering leadership to define the technical roadmap and long-term architecture for serving workloads.
  • •Drive architectural decisions and trade-offs to optimize performance, throughput, autoscaling, and operational efficiency for CPU and GPU serving workloads.
  • •Contribute directly to key components across the serving infrastructure—from model container builds and deployment workflows to runtime systems like routing, caching, observability, and intelligent autoscaling.
  • •Collaborate cross-functionally with product, platform, and research teams to translate customer needs into reliable and performant systems.

Key Requirements

  • •10+ years of experience building and operating large-scale distributed systems.
  • •Deep expertise in model serving, inference systems, and related infrastructure (e.g., routing, scheduling, autoscaling, and observability).
  • •Strong foundation in algorithms, data structures, and system design as applied to large-scale, low-latency serving systems.
  • •Experience leading architecture for large-scale, performance-sensitive CPU/GPU inference systems.
  • •Strong communication skills and ability to collaborate across teams in fast-moving environments.
Experience:10+ yearsData analyticsAI platformCloud
Skills:CommunicationLeadershipProblem-solvingMentoringTeam collaboration
Languages:English
Tech Stack:RoutingAutoscalingObservabilityGPUCPUInferenceDistributed systemsModel serving

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn