Staff Software Engineer, AI Runtime

Databricks
Mountain View, San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 10+ yearsEducation: bachelorsSkills: ["Communication","Mentoring","Collaboration","Leadership","Problem-solving"]

Build and scale Databricks AI Runtime (AIR) across large GPU training fleets, driving architecture, reliability, and developer experience for multi-node training. Lead end-to-end engineering, mentor engineers, and partner with product/research/platform teams to extend AIR, add accelerators, and expand regional coverage while delivering high performance and strong SLAs/SLOs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
2 months ago

Staff Software Engineer, AI Runtime

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Build and scale Databricks AI Runtime (AIR) across large GPU training fleets, driving architecture, reliability, and developer experience for multi-node training. Lead end-to-end engineering, mentor engineers, and partner with product/research/platform teams to extend AIR, add accelerators, and expand regional coverage while delivering high performance and strong SLAs/SLOs.
Location: Mountain View, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Manager level

Key Responsibilities

  • •Lead the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets that span thousands of accelerators.
  • •Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs.
  • •Push GPU efficiency and training performance, raising utilization and lowering cost per training run across diverse model architectures and hardware generations.
  • •Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers.
  • •Partner with product, research, and platform teams to shape APIs, CLI, and developer experience to launch, monitor, and debug production training jobs.

Key Requirements

  • •10+ years of experience building and operating large-scale distributed systems, with significant depth in GPU training infrastructure, high-performance computing, or ML systems.
  • •Hands-on experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models.
  • •Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.
  • •Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (NVLink, InfiniBand or RoCE), collective communication, and bottlenecks that govern training throughput and utilization.
  • •Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability.
Experience:10+ yearsDistributed systemsGPU trainingML systemsCloud
Education:Bachelor's
Skills:CommunicationMentoringCollaborationLeadershipProblem-solving
Languages:English
Tech Stack:PyTorchFSDPDeepSpeedMegatronNVLinkInfiniBandRoCEGPUCUDA

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn