Senior Software Engineer, Agentic AI and Observability

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 168,000 - 270,250 annuallyFunction: Software EngineeringExperience: 8+ yearsEducation: mastersSkills: ["Problem-solving","Debugging","Incident response","Root cause analysis","Cross-team collaboration"]

Build observability and reliability solutions for agentic AI applications in production, covering metrics, tracing, logging, and alerting. Develop deployment pipelines and reliability tooling for the Agentic AI Factory model, define SLOs/SLIs with AI teams, and instrument LLM workflows for performance, cost, and quality monitoring. Own incident response, root cause analysis, and drive observability improvements across the BizApps AI lifecycle.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Software Engineer, Agentic AI and Observability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build observability and reliability solutions for agentic AI applications in production, covering metrics, tracing, logging, and alerting. Develop deployment pipelines and reliability tooling for the Agentic AI Factory model, define SLOs/SLIs with AI teams, and instrument LLM workflows for performance, cost, and quality monitoring. Own incident response, root cause analysis, and drive observability improvements across the BizApps AI lifecycle.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build and implement observability for agentic AI applications in production, including metrics, tracing, logging, and alerting.
  • •Develop and sustain deployment pipelines and reliability tools for the Agentic AI Factory model.
  • •Partner with AI application teams to define SLOs/SLIs and ensure production readiness for new agent deployments.
  • •Instrument LLM-based workflows to monitor performance, cost, and quality, including token usage and tool orchestration.
  • •Drive incident response, root cause analysis, and reliability improvements, closing observability gaps such as hallucination detection and orchestration failure tracing.

Pay and Benefits

Salary: USD 168,000 - 270,250 annually
Equity and Bonus:Equity

Key Requirements

  • •BS or MS in Computer Science, Software Engineering, or a related field (or equivalent experience).
  • •8+ years of experience with strong software engineering skills in Python and modern CI/CD practices.
  • •Hands-on experience with container orchestration (Kubernetes, Docker) and cloud infrastructure.
  • •Experience with observability/monitoring tools such as Datadog, OpenTelemetry, Grafana, or Prometheus.
  • •An SRE/DevOps approach, including ownership of production systems and automating toil for reliability.
Experience:8+ yearsAgentic AIObservabilitySRE/DevOpsDistributed systemsLLM/AI systems in production
Education:Master's
Skills:Problem-solvingDebuggingIncident responseRoot cause analysisCross-team collaboration
Tech Stack:PythonCI/CDKubernetesDockerDatadogOpenTelemetryGrafanaPrometheusLLM orchestrationLangChainLlamaIndexSemantic KernelTerraformPulumiGitOpsContainer orchestrationLLM observability

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor