Manager, Performance Research and Analysis

NVIDIA
Israel
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Experience: 8+ yearsSkills: ["Cross-team leadership","Analytical thinking","Communication skills"]

Lead end-to-end performance research and execution for next-generation NVIDIA AI GPU clusters powering distributed training and inference. Drive characterization, test plans, and optimizations across RDMA/PRDMA networking, collective communication (NCCL), congestion control, and load balancing. Own performance observability by building telemetry-driven dashboards and automated analytics across NICs, switches, GPUs, and NVLink, and perform deep root-cause analysis to mitigate multi-node bottlenecks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Manager, Performance Research and Analysis

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead end-to-end performance research and execution for next-generation NVIDIA AI GPU clusters powering distributed training and inference. Drive characterization, test plans, and optimizations across RDMA/PRDMA networking, collective communication (NCCL), congestion control, and load balancing. Own performance observability by building telemetry-driven dashboards and automated analytics across NICs, switches, GPUs, and NVLink, and perform deep root-cause analysis to mitigate multi-node bottlenecks.
Location: Israel
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Manager level

Key Responsibilities

  • •Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation AI GPU clusters running distributed training and inference workloads.
  • •Evaluate and optimize NVIDIA networking technologies performance, including RDMA/PRDMA, networking protocols, NCCL, congestion control, and load-balancing algorithms.
  • •Conduct performance research on NVIDIA DPUs and storage technologies for North-South use cases supporting AI inference.
  • •Lead performance observability and dashboards for cluster-level analysis using scalable telemetry across NICs, switches, GPUs, and NVLink boundaries.
  • •Perform deep root-cause analysis on complex multi-node performance bottlenecks and drive actionable mitigation plans across hardware, firmware, and software teams.

Key Requirements

  • •B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.
  • •8+ years of experience with deep expertise in High Performance Networking, RDMA, and systems-level performance.
  • •3+ years as an engineering team manager leading technical performance or R&D teams.
  • •Hands-on expertise optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads.
  • •Hands-on experience designing and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
Experience:8+ yearsHigh Performance NetworkingRDMAHPCAIDistributed trainingDistributed inference
Education:
Skills:Cross-team leadershipAnalytical thinkingCommunication skills
Tech Stack:GPUNICSwitchDPUNetworkingRDMAPRDMARoCEv2NCCLMPICollective communicationCongestion controlPFCECNLoad balancingTelemetry pipelinesGrafanaPromQLLogQLNVLink

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor