Senior Software Engineer, AI Runtime

Databricks
Mountain View, San Francisco
Workplace: OnsiteFull timeUSD 160,000 - 225,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Mentoring","Leadership","Collaboration","Problem-solving"]

Senior Software Engineer for AI Runtime at Databricks, leading the architecture and evolution of AIR’s managed GPU training stack, driving high-throughput, multi-node training, and developer experience improvements. You will mentor engineers, collaborate across product, research, and platform teams, and help scale large-scale GPU training infrastructure across thousands of accelerators.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
2 months ago

Senior Software Engineer, AI Runtime

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Senior Software Engineer for AI Runtime at Databricks, leading the architecture and evolution of AIR’s managed GPU training stack, driving high-throughput, multi-node training, and developer experience improvements. You will mentor engineers, collaborate across product, research, and platform teams, and help scale large-scale GPU training infrastructure across thousands of accelerators.
Location: Mountain View, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets that span thousands of accelerators.
  • •Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs.
  • •Push GPU efficiency and training performance, raising utilization and end-to-end throughput while lowering cost per training run across diverse model architectures and hardware generations.
  • •Build resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal customer disruption.
  • •Partner with product, research, and platform teams to shape APIs, CLI, and developer experience to launch, monitor, and debug production training jobs.

Pay and Benefits

Salary: USD 160,000 - 225,000 annually

Key Requirements

  • •5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems.
  • •Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models.
  • •Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.
  • •Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (NVLink, InfiniBand or RoCE), collective communication, and bottlenecks for training throughput and utilization.
  • •Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability.
Experience:5+ yearsAIGPU trainingDistributed systems
Education:Bachelor's
Skills:CommunicationMentoringLeadershipCollaborationProblem-solving
Languages:English
Tech Stack:PyTorchDeepSpeedMegatronFSDPNVLinkInfiniBandRoCEGPUPython

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn