Senior DL Performance Efficiency Architect

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: mastersSkills: ["Technical leadership","Hands-on engineering","Cross-functional collaboration","Measurement-driven approach","Problem solving"]

Lead cross-layer efforts to improve the efficiency of large language models from architecture through training and inference, using measurement-driven analysis to map LLM workloads onto GPUs, memory systems, interconnects, and distributed infrastructure. Drive an efficiency roadmap from early investigation to production deployment, partnering with model researchers, systems and compiler/kernel developers, and hardware architects to co-design future model, software, and computing platforms under practical compute, power, latency, and cost constraints.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior DL Performance Efficiency Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Lead cross-layer efforts to improve the efficiency of large language models from architecture through training and inference, using measurement-driven analysis to map LLM workloads onto GPUs, memory systems, interconnects, and distributed infrastructure. Drive an efficiency roadmap from early investigation to production deployment, partnering with model researchers, systems and compiler/kernel developers, and hardware architects to co-design future model, software, and computing platforms under practical compute, power, latency, and cost constraints.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead cross-layer efforts to improve LLM efficiency across model architecture, training, and inference systems.
  • •Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure to find model-system-hardware co-design opportunities.
  • •Create a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment.
  • •Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence model, software, and hardware roadmaps.
  • •Drive scalable, real-world improvements by identifying fundamental bottlenecks and optimizing within constraints of compute, memory, power, and cost.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •MS or PhD (or equivalent) in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • •5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
  • •Strong understanding of LLM architectures and training/inference workloads, including tradeoffs across quality, compute cost, memory, latency, throughput, and power.
  • •Strong background in performance analysis, roofline modeling, workload characterization, benchmarking, and hardware-aware optimization.
  • •Proven technical leadership to drive complex optimization projects from concept through production deployment.
Experience:5+ yearsAI systemsLarge language modelsHigh-performance computingPerformance optimization
Education:Master's
Skills:Technical leadershipHands-on engineeringCross-functional collaborationMeasurement-driven approachProblem solving
Tech Stack:LLM architecturesGPUsMemory systemsInterconnectsDistributed infrastructurePerformance analysisRoofline modelingBenchmarkingQuantizationSparsityMixture-of-ExpertsLong-context inferenceSpeculative decodingCompilerKernelHardware-aware optimization

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor