Staff Software Engineer - AI Research Infrastructure

Databricks
New York, San Francisco
Workplace: OnsiteFull timeUSD 199,000 - 270,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Distributed systems","GPU","Kubernetes","Slurm","Ray","Rust","C++","Go","Java","Scala","Data pipelines","Model training","Inference workflows"]

Lead design and development of the AI research infrastructure powering Databricks AI Research. Build scalable services to schedule, orchestrate and observe large-scale GPU workloads, improve dev tooling, and ensure researchers can iterate quickly without sacrificing reliability, efficiency, or security. Partner with researchers, ML engineers, and platform teams to turn experimental workloads into robust pipelines and push the limits of our infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
4 months ago

Staff Software Engineer - AI Research Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Lead design and development of the AI research infrastructure powering Databricks AI Research. Build scalable services to schedule, orchestrate and observe large-scale GPU workloads, improve dev tooling, and ensure researchers can iterate quickly without sacrificing reliability, efficiency, or security. Partner with researchers, ML engineers, and platform teams to turn experimental workloads into robust pipelines and push the limits of our infrastructure.
Location: New York, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and implement infrastructure that supports large-scale experiments, data processing, and model training (e.g., HPC clusters, GPU fleets, or cloud-based systems)
  • •Enable researchers to go from idea to large-scale experiment in minutes, not days, by building abstractions for job submission, scheduling, and monitoring
  • •Create tooling that improves research developer productivity, such as experiment management systems and CI/testing infrastructure for research code
  • •Influence the long-term roadmap for research computation, shaping how Databricks AI Research trains, evaluates, and ships models
  • •Serve as a technical mentor for other engineers working on compute, infra, and AI systems

Pay and Benefits

Salary: USD 199,000 - 270,000 annually

Key Requirements

  • •BS/MS or PhD in Computer Science or related field
  • •5+ years of software engineering experience, including substantial time working on large-scale distributed systems or infrastructure
  • •Deep experience building and operating distributed systems, data pipelines, or large-scale backend services, ideally involving GPUs, clusters, or major cloud providers
  • •Proficiency in one or more systems programming languages (e.g., C++, Rust, Go, Java, Scala) and ability to design, implement, and debug complex services
  • •Experience building or contributing to cluster schedulers, resource managers, or large-scale job orchestration systems (e.g., Kubernetes, Slurm, Ray, internal systems)
Experience:5+ yearsDistributed systemsGPUClustersHPCCloud
Skills:Distributed systemsGPUKubernetesSlurmRayRustC++GoJavaScalaData pipelinesModel trainingInference workflows
Languages:English
Tech Stack:C++RustGoJavaScalaKubernetesSlurmRayGPUHPCCloudDistributed systems

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn