Senior Solutions Architect, Generative AI

NVIDIA
United States
Workplace: RemoteFull timeUSD 184,000 - 356,500 annuallyFunction: Solutions Engineering & Sales EngineeringEducation: bachelorsSkills: ["Collaboration","Technical problem-solving","Performance analysis"]

Work with customers and frontier labs to accelerate end-to-end AI workloads on large-scale GPU systems. Design and optimize high-performance AI clusters spanning compute, networking, storage, scheduling, orchestration, and observability. Profile training and inference to find bottlenecks, diagnose distributed systems issues across InfiniBand/RoCE and RDMA stacks, and lead POCs and performance studies. Partner across engineering, product, and sales to win design approvals and deliver technical solutions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior Solutions Architect, Generative AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Work with customers and frontier labs to accelerate end-to-end AI workloads on large-scale GPU systems. Design and optimize high-performance AI clusters spanning compute, networking, storage, scheduling, orchestration, and observability. Profile training and inference to find bottlenecks, diagnose distributed systems issues across InfiniBand/RoCE and RDMA stacks, and lead POCs and performance studies. Partner across engineering, product, and sales to win design approvals and deliver technical solutions.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Collaborate with customers to maximize GPU utilization and improve end-to-end workload throughput while enhancing reliability and reducing infrastructure costs.
  • •Design and optimize large-scale AI clusters across GPU compute, networking, storage, workload scheduling, orchestration, and observability.
  • •Profile distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, fabrics, storage systems, and software stack.
  • •Diagnose complex infrastructure and distributed systems issues across InfiniBand/RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
  • •Lead proof-of-concepts and performance studies, creating benchmarking tools, automation, runbooks, and technical collateral; partner with engineering, product, and sales to drive design wins.
Travel: Medium travel

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.
  • •BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another engineering field (or equivalent experience).
  • •Deep understanding of Linux systems, distributed computing, GPU architectures, and large-scale AI cluster hardware/software.
  • •Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks using InfiniBand, RoCE, or GPUDirect RDMA.
  • •Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.
Experience:AI infrastructureHigh-performance computingGPU systemsDistributed computingNetworkingSite reliability engineering
Education:Bachelor's
Skills:CollaborationTechnical problem-solvingPerformance analysis
Tech Stack:LinuxInfiniBandRoCEGPUDirect RDMARDMANCCLNVLinkNVSwitchDGXHGXSpectrum-XDCGMNsight SystemsKubernetesSlurmPythonShell scriptingContainersObservability

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor