Sr. Engineering Manager, AI Runtime

Databricks
Mountain View, San Francisco
Workplace: OnsiteFull timeUSD 228,600 - 297,120Function: Data Science & Machine LearningExperience: 8+ yearsEducation: bachelorsSkills: ["Leadership","Mentorship","Cross-functional collaboration","Communication","Decision-making under ambiguity"]

Lead and grow the engineering team behind AI Runtime’s Custom Training product, owning both customer-facing experience and foundational infrastructure for managed GPU training. Define and drive the AIR product and technical roadmap, orchestrate distributed training, and improve cluster lifecycle, fault tolerance, and training efficiency. Partner across platform, product, research, and customers to deliver and operate resilient, scalable multi-node training systems with strong observability and reliability practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
2 months ago

Sr. Engineering Manager, AI Runtime

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead and grow the engineering team behind AI Runtime’s Custom Training product, owning both customer-facing experience and foundational infrastructure for managed GPU training. Define and drive the AIR product and technical roadmap, orchestrate distributed training, and improve cluster lifecycle, fault tolerance, and training efficiency. Partner across platform, product, research, and customers to deliver and operate resilient, scalable multi-node training systems with strong observability and reliability practices.
Location: Mountain View, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Manager level

Key Responsibilities

  • •Lead, mentor, and grow a high-performing engineering team for AIR Custom Training, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.
  • •Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • •Collaborate with product, research, platform, infrastructure teams, and customers to deliver end-to-end from ideation and prioritization to launch and operation.
  • •Drive architectural decisions and product design for managed GPU training at scale.
  • •Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.

Pay and Benefits

Salary: USD 228,600 - 297,120
Equity and Bonus:Equity

Key Requirements

  • •8+ years of software engineering experience, including 3+ years in engineering management.
  • •Track record building and operating managed GPU training infrastructure at scale (100s/1000s GPUs).
  • •Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism).
  • •Experience with training resilience patterns such as checkpointing, elastic training, and automated failure recovery for long-running jobs.
  • •BS/MS in Computer Science, Electrical Engineering, or a related technical field.
Experience:8+ yearsAIMachine learningDeep learningLLMGPU trainingDistributed training
Education:Bachelor's in Computer Science or Electrical Engineering
Skills:LeadershipMentorshipCross-functional collaborationCommunicationDecision-making under ambiguity
Languages:English
Tech Stack:PyTorchDeepSpeedComposerMegatron-LMFSDPTensor parallelismPipeline parallelismNCCLLLMGPUsDistributed training orchestration

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn