Senior Software Architect, AI Networking

NVIDIA
Tel Aviv
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","System-level thinking"]

Shape scalable architectures for multi-node LLM inference across GPU clusters, optimizing latency, throughput, and cost. Collaborate across model, systems, compiler, and networking teams to deliver high-performance solutions, including memory orchestration, compute scheduling, inter-node communication, and system-level optimizations. Prototype KV cache handling, tensor/pipeline parallel execution, and dynamic batching, and translate architecture into production systems through design docs, specs, and technical writing.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 days ago

Senior Software Architect, AI Networking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live
Reposted: similar role first listed 8 months ago

Job Summary

Shape scalable architectures for multi-node LLM inference across GPU clusters, optimizing latency, throughput, and cost. Collaborate across model, systems, compiler, and networking teams to deliver high-performance solutions, including memory orchestration, compute scheduling, inter-node communication, and system-level optimizations. Prototype KV cache handling, tensor/pipeline parallel execution, and dynamic batching, and translate architecture into production systems through design docs, specs, and technical writing.
Location: Tel Aviv
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and evolve scalable architectures for multi-node LLM inference across GPU clusters.
  • •Develop infrastructure to optimize latency, throughput, and cost-efficiency of serving large models in production.
  • •Collaborate across model, systems, compiler, and networking teams to ensure complete, high-performance solutions.
  • •Prototype novel approaches to KV cache handling, tensor/pipeline parallel execution, and dynamic batching.
  • •Author design documents and internal specs, and contribute to open-source efforts when appropriate.

Key Requirements

  • •Bachelor’s, Master’s, or PhD in Computer Science, Electrical Engineering, or equivalent experience.
  • •8+ years building large-scale distributed systems or performance-critical software.
  • •Deep understanding of deep learning systems, GPU acceleration, and AI model execution flows and/or high performance networking.
  • •Proven software engineering skills in C++ and/or Python, with strong familiarity with CUDA or similar platforms.
  • •Strong system-level thinking across memory, networking, scheduling, and compute orchestration.
Experience:8+ yearsDistributed systemsAIDeep learningGPU acceleration
Education:Bachelor's in Computer Science, Electrical Engineering
Skills:CommunicationCollaborationSystem-level thinking
Tech Stack:C++PythonCUDA

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor