Software Engineer - GPU Networking & Distributed Systems

Baseten
San Francisco, Toronto, Canada, New York, Montreal
Workplace: HybridFull timeUSD 150,000 - 250,000 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Communication","Collaboration","Attention to detail"]

Foundational software engineer focused on GPU networking and distributed systems to co-optimize communication and compute for disaggregated serving, WideEP, and low-latency inference across multi-GPU clusters. You will work on RDMA integration, KV cache offloads, and hardware-focused performance, leveraging NCCL/NVSHMEM, TensorRT-LLM, and NVLink in a hybrid infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
6 months ago

Software Engineer - GPU Networking & Distributed Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Foundational software engineer focused on GPU networking and distributed systems to co-optimize communication and compute for disaggregated serving, WideEP, and low-latency inference across multi-GPU clusters. You will work on RDMA integration, KV cache offloads, and hardware-focused performance, leveraging NCCL/NVSHMEM, TensorRT-LLM, and NVLink in a hybrid infrastructure.
Location: San Francisco, Toronto, Canada, New York, Montreal
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Make RDMA First-Class by integrating RDMA/RoCE/InfiniBand into the inference stack to improve bandwidth and latency
  • •Optimize distributed inference by tuning networking layers for Disaggregated KV Cache Offload and WideEP across NVLink and InfiniBand
  • •Enable sub-10-second startup for trillion-parameter models by advancing checkpointing and storage mechanisms
  • •Deep-dive into hardware performance on cutting-edge clusters (H100/H200, NVL72) and write acceptance tests for peak throughput and low latency
  • •Build observability tools to visualize packet flow, congestion, and bandwidth across GPU interconnects

Pay and Benefits

Salary: USD 150,000 - 250,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kPaid LeaveEquity

Key Requirements

  • •Deep experience with high-performance networking protocols (InfiniBand, RoCE v2) and understanding of data movement physics
  • •Fluent in C++ or Python with ability to bridge high-level logic and hardware
  • •Strong understanding of memory hierarchy in modern NVIDIA architectures (H100/Blackwell) and optimization for it
  • •Willingness to dive into TensorRT-LLM source code and write custom C++/Python bindings
  • •Experience with or willingness to work with NCCL, NVSHMEM, and UCX
Experience:High Performance ComputingAI infrastructureDistributed systems
Skills:Problem-solvingCommunicationCollaborationAttention to detail
Languages:English
Tech Stack:C++PythonNCCLNVSHMEMUCXInfiniBandRoCE v2NVLinkTensorRT-LLMGPUH100B100NVL72Kubernetes

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor