Senior Staff Site Reliability Engineer

NVIDIA
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 15+ yearsEducation: bachelorsSkills: ["Analytical skills","Automation","Monitoring","Capacity planning","Collaboration"]

Design, scale, and deploy core IT infrastructure services for global on-prem and cloud environments. Lead initiatives for service architecture across DNS, NTP/PTP, DHCP, and LDAP, with a focus on performance, reliability, automation, monitoring, and lifecycle management. Define efficiency metrics, build tools for data analysis and alerting, and optimize observability and DDoS mitigation using technologies like eBPF and XDP. Collaborate across leadership to deliver IT products and services.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Design, scale, and deploy core IT infrastructure services for global on-prem and cloud environments. Lead initiatives for service architecture across DNS, NTP/PTP, DHCP, and LDAP, with a focus on performance, reliability, automation, monitoring, and lifecycle management. Define efficiency metrics, build tools for data analysis and alerting, and optimize observability and DDoS mitigation using technologies like eBPF and XDP. Collaborate across leadership to deliver IT products and services.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead initiatives to transform the IT Compute Core Team architecture and drive new service offerings across on-prem and cloud.
  • •Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP with performance and reliability at global scale.
  • •Define and implement metrics for service efficiency and drive improvements via software/hardware optimizations (SR-IOV/DPU).
  • •Build and maintain tools for collecting, analyzing, and visualizing data to support reporting, alerting, and monitoring.
  • •Collaborate with leadership, senior engineers, program managers, and product managers to develop IT products and services that meet customer needs.

Key Requirements

  • •Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience.
  • •15+ years of proven experience in compute platform engineering focused on automation.
  • •Experience designing and deploying containerization architectures and distributed systems infrastructure.
  • •Proficiency with programming languages such as Go and/or Python and strong Linux/kernel fundamentals.
  • •Experience developing tools for data analysis/performance profiling and with Terraform and configuration management tools.
Experience:15+ yearsCompute platform engineeringInfrastructureOn-premCloudContainerizationDistributed systemsMicroservices
Education:Bachelor's
Skills:Analytical skillsAutomationMonitoringCapacity planningCollaboration
Tech Stack:DNSNTPPTPDHCPLDAPEBPFXDPSR-IOVDPUTerraformContainerizationDistributed systemsMicroservicesInfrastructure as code (IaC)Configuration managementGoPythonLinuxLinux kernel internalsVLAN

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor