Senior Deep Learning Systems Engineer, Datacenters

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: IT Operations (Systems/Network Admin)Experience: 8+ yearsEducation: bachelorsSkills: ["Python","C++","TensorFlow","PyTorch","Linux","CUDA","Docker","Slurm","Performance profiling","Gprof","Nvidia-smi","Dcgm","Bash"]

Lead performance analysis and optimization of datacenter hardware and software for AI workloads. Design and evaluate DL deployments across CPU/GPU, memory, and networking, building tools to characterize, profile, and improve efficiency of DL workloads on NVIDIA systems across hybrid datacenter environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Senior Deep Learning Systems Engineer, Datacenters

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead performance analysis and optimization of datacenter hardware and software for AI workloads. Design and evaluate DL deployments across CPU/GPU, memory, and networking, building tools to characterize, profile, and improve efficiency of DL workloads on NVIDIA systems across hybrid datacenter environments.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Sr. Manager level

Key Responsibilities

  • •Develop software infrastructure to characterize and analyze a broad range of Deep Learning applications.
  • •Evolve cost-efficient datacenter architectures tailored to meet the needs of large language models (LLMs).
  • •Work with experts to develop analysis and profiling tools in Python, bash and C++ to measure key performance metrics of DL workloads on NVIDIA systems.
  • •Analyze system and software characteristics of DL applications.
  • •Develop analysis tools and methodologies to measure key performance metrics and to estimate potential for efficiency improvement.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity
Perks:EquityBenefits401kHealth InsuranceRemote WorkPaid Leave

Key Requirements

  • •A Bachelor’s degree in Electrical Engineering or Computer Science or equivalent experience (Masters or PhD degree preferred).
  • •8 years or more of relevant experience.
  • •Experience in System Software: Linux OS, compilers, CUDA kernels, DL frameworks (PyTorch, TensorFlow) OR Silicon Architecture and Performance Modeling/Analysis (CPU, GPU, memory, or network).
  • •Experience programming in C/C++ and Python; exposure to Docker and Slurm is a plus.
  • •Strong understanding of computer system architecture and performance analysis; ability to own tasks from start to finish; experience in virtual environments.
Experience:8+ yearsDatacenterAIDeep learningDL
Education:Bachelor's
Skills:PythonC++TensorFlowPyTorchLinuxCUDADockerSlurmPerformance profilingGprofNvidia-smiDcgmBash
Tech Stack:PythonC++CUDALinuxPyTorchTensorFlowDockerSlurmNvidia-smiDcgmGprofBash

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor