Senior AI infrastructure engineer - EDA Infrastructure

NVIDIA
California, Austin, Washington, Durham
Workplace: RemoteFull timeUSD 184,000 - 356,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Cross-functional leadership","Proactive risk mitigation","Problem-solving"]

Build and operate scalable telemetry and observability infrastructure for NVIDIA GPU cloud services, including pipelines for metrics, logs, traces, and events. Standardize and automate incident and on-call workflows, improve operational responsiveness with reporting and AI-assisted tooling, and maintain trusted hardware/software catalogs and service ownership data. Partner across engineering and external teams to deliver production-ready platforms with strong instrumentation and lifecycle signals.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Senior AI infrastructure engineer - EDA Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and operate scalable telemetry and observability infrastructure for NVIDIA GPU cloud services, including pipelines for metrics, logs, traces, and events. Standardize and automate incident and on-call workflows, improve operational responsiveness with reporting and AI-assisted tooling, and maintain trusted hardware/software catalogs and service ownership data. Partner across engineering and external teams to deliver production-ready platforms with strong instrumentation and lifecycle signals.
Location: California, Austin, Washington, Durham
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and operate telemetry pipelines for metrics, logs, traces, and events across on-premise, CSP, and NCP clusters.
  • •Establish standard instrumentation, collection, storage, and access patterns for consistent telemetry.
  • •Deliver dashboards, alerting, and analysis to improve visibility, detection, and troubleshooting.
  • •Standardize and automate incident, maintenance, and on-call workflows across HWInf, including integration of operational data and lifecycle signals.
  • •Build and maintain hardware/software catalogs as sources of truth for infrastructure inventory, service ownership, dependencies, and documentation.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS degree in Computer Science, Computer Engineering, or related field, or equivalent experience.
  • •8+ years of experience in infrastructure security, platform engineering, or security tooling.
  • •Proficiency in one or more programming languages such as Python, Go, TypeScript, or Java.
  • •Strong understanding of software and infrastructure principles with production experience.
  • •Ability to lead cross-functional initiatives across internal teams and external partners across engineering, product, finance, and security.
Experience:8+ yearsInfrastructure securityPlatform engineeringObservabilityAI/ML in productionSaaS
Education:Bachelor's
Skills:Cross-functional leadershipProactive risk mitigationProblem-solving
Tech Stack:PythonGoTypeScriptJavaSaaSML modelsObservabilityTelemetryMetricsLogsTracesProfilingCMDB

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor