Principal Engineer, Compute Fleet Management

Databricks
Bellevue, United States
Workplace: OnsiteFull timeUSD 264,300 - 322,300 annuallyFunction: Transportation & Fleet OperationsSkills: ["Leadership","Influence","Execution","Cross-team collaboration","Problem-solving"]

Lead the compute fleet management at Databricks, owning the architecture and optimization of cross-cloud compute resources (AWS, Azure, GCP) to achieve peak workload performance, high availability, and strong isolation. You will drive large, cross-team infrastructure initiatives, scale distributed systems, and influence technical direction across the organization to improve margin and customer experience.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
8 months ago

Principal Engineer, Compute Fleet Management

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead the compute fleet management at Databricks, owning the architecture and optimization of cross-cloud compute resources (AWS, Azure, GCP) to achieve peak workload performance, high availability, and strong isolation. You will drive large, cross-team infrastructure initiatives, scale distributed systems, and influence technical direction across the organization to improve margin and customer experience.
Location: Bellevue, United States
Workplace: Onsite
Employment Type: Full time
Job Function: Transportation & Fleet Operations

Key Responsibilities

  • •Pioneer fleet optimization by provisioning and pooling of billions of cloud resources to achieve peak workload performance, efficiency, and robust resource isolation.
  • •Deliver hyper-scale resilience by building architecture that guarantees horizontal scaling and resilience against cloud/account-level failures.
  • •Own the critical path by leading the development of the lowest-dependency systems required to bootstrap and manage the massive compute platform.
  • •Ensure high availability by achieving and maintaining 99.99% availability for all batch and serving workloads.
  • •Drive execution discipline with planning, tracking, and managing complex cross-organizational dependencies.

Pay and Benefits

Salary: USD 264,300 - 322,300 annually
Equity and Bonus:Equity

Key Requirements

  • •Leading transformative projects: taking ownership of complex, cross-team, cross-layer, and multi-quarter strategic engineering initiatives from concept to execution.
  • •Distributed systems mastery: deep, hands-on experience developing and operating high-scale distributed systems on at least one major public cloud.
  • •Influence without authority: proven ability to drive consensus, establish technical direction, and lead large technical efforts across organizational boundaries.
  • •Execution discipline: exceptional strength in planning, tracking project progress, and managing complex cross-organizational dependencies.
  • •Experience managing and scaling a massive fleet of GPUs for AI/ML workloads and operating large-scale distributed systems across all major clouds (AWS, Azure, and GCP).
Experience:Cloud computingDistributed systemsAI/MLInfrastructure
Skills:LeadershipInfluenceExecutionCross-team collaborationProblem-solving
Languages:English
Tech Stack:AWSAzureGCPDistributed systemsGPU computingCloud infrastructure

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn