Principal Networking AI Systems Architect

NVIDIA
Tel Aviv
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 10+ yearsEducation: phdSkills: ["Communication","Collaboration"]

Scope and lead AI-based solutions for data-center networking management, including agentic AI for troubleshooting, predictive resiliency/AIOPS, black-box optimization, and performance tuning. Partner with engineering and reliability teams to define AI/ML workflows and integrate AI capabilities into system architecture and engineering processes. Identify opportunities for automated failure management, performance improvement, and resource optimization by translating system constraints and dependencies into research problems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Principal Networking AI Systems Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Scope and lead AI-based solutions for data-center networking management, including agentic AI for troubleshooting, predictive resiliency/AIOPS, black-box optimization, and performance tuning. Partner with engineering and reliability teams to define AI/ML workflows and integrate AI capabilities into system architecture and engineering processes. Identify opportunities for automated failure management, performance improvement, and resource optimization by translating system constraints and dependencies into research problems.
Location: Tel Aviv
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Build a shared roadmap and vision for AI-based data-center management solutions spanning LLM intelligence, predictive resiliency/AIOPS, black-box optimization, and performance tuning.
  • •Collaborate with engineering and reliability teams to scope and define AI/ML-driven workflows.
  • •Drive the integration of AI capabilities into system architecture and engineering workflows.
  • •Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization.
  • •Translate system behavior, dependencies, data, and operational constraints into formulated research problems.

Key Requirements

  • •Ph.D. in electrical engineering, machine learning, computer science, or a related field.
  • •10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design.
  • •Expertise in high-speed interconnect protocols, including InfiniBand and/or advanced Ethernet architectures.
  • •Proven experience driving high-impact AI/ML projects, including LLMs/agents, deep learning, and black-box optimization.
  • •Strong knowledge of AI/ML and networking hardware/system architecture, with the ability to communicate data-based insights to stakeholders.
Experience:10+ years
Education:PhD / Doctorate
Skills:CommunicationCollaboration
Tech Stack:AI/MLLLMsAgentsDeep learningBlack-box optimizationAIOPSData-center managementInfiniBandEthernetBlueField DPUsQuantum InfiniBand switchesSpectrum EthernetIn-network computingTelemetryAdaptive routingTelemetry-driven network optimizationGPU clustersDistributed systems design

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor