Senior Software Developer - Network and Collectives

Intel
Israel
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Concurrency","Lock-free design","Performance engineering","Profiling","Numerical determinism"]

Build and optimize topology-aware collective communication algorithms for Intel’s AI accelerator hardware fabric. Own the design and implementation of collective operations (all-reduce, all-gather, reduce-scatter, point-to-point) tuned to interconnect bandwidth/latency, and develop the transport layer across tray/rack/pod. Profile multi-node performance to close the gap to fabric peak, and co-design with hardware and multi-node runtime teams for partitioning, overlap, and scheduling while ensuring correctness and numerical determinism.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Intel
Intel
1 month ago

Senior Software Developer - Network and Collectives

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize topology-aware collective communication algorithms for Intel’s AI accelerator hardware fabric. Own the design and implementation of collective operations (all-reduce, all-gather, reduce-scatter, point-to-point) tuned to interconnect bandwidth/latency, and develop the transport layer across tray/rack/pod. Profile multi-node performance to close the gap to fabric peak, and co-design with hardware and multi-node runtime teams for partitioning, overlap, and scheduling while ensuring correctness and numerical determinism.
Location: Israel
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and implement collective algorithms tuned to interconnect topology and bandwidth/latency profile.
  • •Build a topology-aware transport layer across tray/rack/pod interconnects.
  • •Optimize end-to-end collective performance across multi-node pods by profiling and eliminating bottlenecks.
  • •Co-design with hardware on interconnect features and with multi-node runtime on partitioning, overlap, and scheduling.
  • •Own correctness and numerical determinism of reductions across scale-out.

Key Requirements

  • •5+ years building AI/Systems/HPC software in C++ with strong concurrency and lock-free design.
  • •Hands-on experience with collective libraries (e.g., NCCL/MPI) or scale-out communication.
  • •Knowledge of communication patterns behind tensor, pipeline, and expert parallelism and how they map to collectives and underlying fabric.
  • •Performance engineering experience including profiling, bandwidth/latency tuning, and roofline reasoning.
  • •Nice to have: experience with RDMA/InfiniBand/RoCE, GPUDirect-style transfers, or comparable fabric technologies.
Experience:5+ yearsAIHPCDistributed training/inferenceScale-out communication
Skills:ConcurrencyLock-free designPerformance engineeringProfilingNumerical determinism
Tech Stack:C++NCCLMPIVLLMSGLangTensorRT-LLMDeepSeekRDMAInfiniBandRoCEGPUDirect

Company Brief

Intel
Designs and manufactures semiconductor chips, processors, and related hardware for PCs, data centers, networking, and embedded applications, while providing software and services to accelerate computing across industries globally.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1968
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor