Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA
Santa Clara, United States
Full timeUSD 152,000 - 241,500 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Ownership","Automation"]

Design, implement, and operate large-scale Observability & Telemetry platform services focused on real-time monitoring, logging, and alerting. Own the end-to-end lifecycle from inception and design through deployment, operation, and continuous refinement, including capacity planning and launch reviews. Maintain high availability and low latency by measuring system health and scaling sustainably via automation, incident response, and blameless postmortems. Participate in an on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Systems Software Engineer, Observability and Telemetry Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design, implement, and operate large-scale Observability & Telemetry platform services focused on real-time monitoring, logging, and alerting. Own the end-to-end lifecycle from inception and design through deployment, operation, and continuous refinement, including capacity planning and launch reviews. Maintain high availability and low latency by measuring system health and scaling sustainably via automation, incident response, and blameless postmortems. Participate in an on-call rotation.
Location: Santa Clara, United States
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and support operational and reliability aspects of the Observability & Telemetry collection platform, with performance at scale.
  • •Engage in the full service lifecycle from inception and design through deployment, operation, and refinement.
  • •Support pre-launch activities including system design consulting, developing tools/platforms/frameworks, capacity management, and launch reviews.
  • •Maintain live services by measuring and monitoring availability, latency, and overall system health.
  • •Scale sustainably through automation, practice incident response and blameless postmortems, and participate in an on-call rotation.

Pay and Benefits

Salary: USD 152,000 - 241,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • •5+ years of experience with infrastructure automation and distributed systems design.
  • •5+ years delivering foundational infrastructure and observability platforms.
  • •Experience with one or more: Python, Go, Perl, or Ruby.
  • •In-depth knowledge of Linux, networking, and containers, including Kubernetes, OpenStack, and Docker.
Experience:5+ yearsInfrastructure automationDistributed systemsObservability platformsPrivate or public cloud
Education:Bachelor's in Computer Science
Skills:Problem-solvingCommunicationOwnershipAutomation
Tech Stack:PythonGoPerlRubyLinuxKubernetesOpenStackDockerGrafanaOpenTelemetryPrometheus

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor