Senior Staff Site Reliability Engineer - Compute Core Engineering

NVIDIA
Santa Clara, United States
Workplace: RemoteFull timeUSD 200,000 - 322,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 12+ yearsSkills: ["Analytical skills","Collaboration","Autonomous"]

Lead IT Compute Core initiatives to transform services across on-prem and cloud, designing and operating scalable infrastructure for global reliability. Build and automate core services (DNS, NTP/PTP, DHCP, LDAP) with monitoring, high availability, capacity planning, and lifecycle management. Define service efficiency metrics, develop data tooling for reporting/alerting, and collaborate with engineering and product leaders to deliver infrastructure products and services that meet customer needs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 days ago

Senior Staff Site Reliability Engineer - Compute Core Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead IT Compute Core initiatives to transform services across on-prem and cloud, designing and operating scalable infrastructure for global reliability. Build and automate core services (DNS, NTP/PTP, DHCP, LDAP) with monitoring, high availability, capacity planning, and lifecycle management. Define service efficiency metrics, develop data tooling for reporting/alerting, and collaborate with engineering and product leaders to deliver infrastructure products and services that meet customer needs.
Location: Santa Clara, United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead initiatives to transform the IT Compute Core Team architecture and build new service offerings across on-prem and cloud.
  • •Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP, focusing on performance and reliability at global scale.
  • •Define and implement metrics to measure service efficiency, driving efficiency improvements through software and hardware optimizations (SR-IOV/DPU).
  • •Collect and analyze system data for capacity planning, develop enterprise-wide plans, and coordinate changes with management.
  • •Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, and monitoring, collaborating with engineering leadership to meet customer needs.

Pay and Benefits

Salary: USD 200,000 - 322,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience.
  • •12+ years of proven experience in compute platform engineering with a focus on automation.
  • •Experience designing, deploying, and evaluating containerization architectures and distributed systems infrastructure to improve scalability, reliability, and efficiency.
  • •Proficiency in programming languages such as Go and/or Python, plus Linux OS proficiency with kernel internals.
  • •Experience with infrastructure automation and data/performance tooling (Terraform, configuration management tools), and working with large bare metal environments.
Experience:12+ yearsCompute platform engineeringAutomationDistributed systemsContainersBare metal infrastructure
Education:
Skills:Analytical skillsCollaborationAutonomous
Tech Stack:DNSNTPPTPDHCPLDAPAutomationMonitoringHigh availabilityCapacity planningLifecycle managementEBPFXDPSR-IOVDPUTerraformConfiguration managementGoPythonLinuxKernel internals

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor