Senior Platform Telemetry Engineer

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 152,000 - 287,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: mastersSkills: ["Communication","Problem-solving","Teamwork","Leadership"]

Senior Platform Telemetry Engineer to design and deliver scalable telemetry, fleet management, and health-monitoring solutions for AI infrastructure. You’ll work across customers, product management, and architects to define requirements, craft architecture specs, drive end-to-end delivery, and ensure robust testing and productization of telemetry and REST-based interfaces on GPU-powered platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Senior Platform Telemetry Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Senior Platform Telemetry Engineer to design and deliver scalable telemetry, fleet management, and health-monitoring solutions for AI infrastructure. You’ll work across customers, product management, and architects to define requirements, craft architecture specs, drive end-to-end delivery, and ensure robust testing and productization of telemetry and REST-based interfaces on GPU-powered platforms.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Drive next generation fleet management solutions for scaling AI infrastructure using GPUs and Grace solution from Nvidia; collaborate with customers, product management and architects to finalize requirements for rapid product development.
  • •Define and clarify architecture for fleet health monitoring and fault remediation at scale; perform POCs, document architecture, and ensure capabilities are used effectively both in-band and out-of-band.
  • •Educate customers on product architecture, incorporate feedback, write architecture specs and design documents, and own end-to-end delivery across teams; perform code reviews related to architecture specs.
  • •Collaborate with development to ensure proper unit testing and test plans; drive QA to productize code and assume product ownership.
  • •Capture requirements in Jira, manage bug tracking, and collaborate with managers to develop end-to-end execution plans; contribute to all phases of product development from definition to early customer support.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Equity and Bonus:Equity
Perks:EquityBenefits

Key Requirements

  • •BS, MS, or PhD in EE/CS or related field (or equivalent experience).
  • •5+ years hands-on coding experience.
  • •Strong knowledge of time series databases like Influxdb & Prometheus; REST APIs (Redfish a plus); telemetry visualization with Grafana & Influx; firmware architecture and low-latency APIs; analysis of time/space complexity and resource requirements.
  • •Proven record of scalability solutions.
  • •Strong C/C++ and Python skills; server platform programming and debugging; experience with SCM (Git, Perforce) and Jira.
Experience:5+ yearsAIHPCTelemetryServersGPU
Education:Master's
Skills:CommunicationProblem-solvingTeamworkLeadership
Tech Stack:CC++PythonREST APIsGrafanaInfluxDBPrometheusRedfishGitPerforceJiraFirmwareARMX86Open Compute

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor