Driver Engineer

Modular
United States, Canada
Workplace: HybridFull timeUSD 167,000 - 242,000 annuallyFunction: Transportation & Fleet OperationsSkills: ["Ownership","Concurrency","Debugging","Collaboration","Problem-solving"]

Build low-level driver infrastructure that sits between the compiler/runtime stack and accelerator silicon, enabling MAX and Mojo to run reliably and efficiently across NVIDIA, AMD, Apple Silicon, and new hardware. You’ll implement core driver abstractions (device, context, queue, memory, kernel launch), develop multi-accelerator/multi-node communication primitives, and improve async diagnostics for kernel and graph authors. Work across kernels, compiler/runtime, and serving on a fully open-source stack.

This position is no longer accepting applications.

  • See live roles at Modular
  • Search all live jobs
  • Browse companies, collections, and locations hiring now
Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa

This position is no longer accepting applications.

See live roles at ModularSearch all live jobsBrowse companies, collections, and locations hiring now

Modular
Modular
3 months ago

Driver Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: Just nowStatus: Closed

Job Summary

Build low-level driver infrastructure that sits between the compiler/runtime stack and accelerator silicon, enabling MAX and Mojo to run reliably and efficiently across NVIDIA, AMD, Apple Silicon, and new hardware. You’ll implement core driver abstractions (device, context, queue, memory, kernel launch), develop multi-accelerator/multi-node communication primitives, and improve async diagnostics for kernel and graph authors. Work across kernels, compiler/runtime, and serving on a fully open-source stack.
Location: United States, Canada
Workplace: Hybrid
Employment Type: Full time
Job Function: Transportation & Fleet Operations
Seniority: Mid level

Key Responsibilities

  • •Implement and extend core driver abstractions (Device, Context, Queue, Memory, Function/kernel representation and launch) across diverse hardware backends.
  • •Build multi-accelerator and multi-node communication and collectives primitives for large-model inference using technologies like NVLink, RDMA, and UCX.
  • •Improve diagnostics and error reporting across an asynchronous execution stack, converting low-level failures into actionable messages for production authors.
  • •Partner with Kernels, Graph Compiler/Runtime, and Serving teams to shape the surfaces used by Mojo standard library and MAX framework.
  • •Participate in design discussions and code reviews while contributing to a fully open-source project with public GitHub contributions.
Travel: Low travel

Pay and Benefits

Salary: USD 167,000 - 242,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid Leave

Key Requirements

  • •3+ years writing high-performance, low-latency production systems in C++ (modern C++17/20), with strong ownership, lifetime, ABI stability, concurrency, and parallelism.
  • •Hands-on experience with at least one accelerator driver-level API such as CUDA Driver API, HIP, Metal, or Vulkan compute (streams, events, contexts, module loading).
  • •Working understanding of accelerator execution and memory models including stream ordering, host↔device transfers, and pinned memory.
  • •Solid instincts for library and API design, including naming, layering, and ergonomics.
  • •Strong debugging skills in asynchronous, multi-device systems using tools like GDB/LLDB and sanitizers.
Skills:OwnershipConcurrencyDebuggingCollaborationProblem-solving
Tech Stack:C++C++17C++20CUDA Driver APIHIPMetalVulkanStreamsEventsContextsModule loadingAsync allocatorsIPCNVLinkRDMAInfinibandRoCEEFASocketsUCX

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Modular
Provides infrastructure and developer tools to build, deploy, and scale large AI models and foundation-model applications, including model hosting, orchestration, and SDKs to accelerate AI product development.
Industry: AI & Machine Learning
Website