Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA
Gurugram
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: mastersSkills: ["Communication","Interpersonal skills","Customer-facing skills","Listening skills","Troubleshooting"]

Build AI/HPC infrastructure and support large-scale AI clusters for customers, focusing on performance, monitoring, logging, and alerting. Own the service lifecycle from inception and design through deployment and refinement, and develop automation tooling for network provisioning, observability, and self-service resource consumption. Troubleshoot from bare metal to application level and contribute to POCs/POVs while documenting reusable methodologies for customer and internal teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Solutions Architect, Networking and Compute Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build AI/HPC infrastructure and support large-scale AI clusters for customers, focusing on performance, monitoring, logging, and alerting. Own the service lifecycle from inception and design through deployment and refinement, and develop automation tooling for network provisioning, observability, and self-service resource consumption. Troubleshoot from bare metal to application level and contribute to POCs/POVs while documenting reusable methodologies for customer and internal teams.
Location: Gurugram
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Build AI/HPC infrastructure for new and existing customers.
  • •Support operational and reliability aspects of large-scale AI clusters, including real-time monitoring, logging, and alerting.
  • •Manage the full lifecycle of services from inception and design through deployment, operation, and refinement.
  • •Develop tooling to automate and manage large-scale infrastructure, including automated provisioning, observability, and self-service consumption.
  • •Deploy monitoring solutions and troubleshoot bottom up from bare metal through OS, software stack, and applications; support POCs/POVs and document methodologies.
Travel: Low travel

Key Requirements

  • •5+ years work or research experience in networking fundamentals (TCP/IP) and data center compute architecture.
  • •Advanced knowledge of HPC and AI networking protocols (EVPN, BGP, OSPF, VXLAN).
  • •Deep understanding of DC architecture fundamentals including compute, storage, InfiniBand, Ethernet, and NVLink.
  • •Python and bash scripting experience with strong Linux systems administration (CentOS, RHEL, Ubuntu).
  • •Experience with automated network provisioning and automation/config management tools (e.g., Jenkins, Ansible, Puppet, Chef).
Experience:5+ yearsHPCAIData centers
Education:Master's
Skills:CommunicationInterpersonal skillsCustomer-facing skillsListening skillsTroubleshooting
Certifications:CCNPCCIELinux certificationsNVIDIA-related certifications
Languages:English
Tech Stack:LinuxCentOSRHELUbuntuPythonBashTCP/IPHPCAIEVPNBGPOSPFVXLANEthernetInfiniBandNVLinkRDMARoCEPFSInfiniBand/RDMA

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor