Senior AI Tools Engineer, SRE Operations - GeForce NOW

NVIDIA
Santa Clara, California
Workplace: RemoteFull timeUSD 144,000 - 230,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: []

Build and deploy AI-powered tools that transform GeForce NOW production signals, metrics, and logs into actionable intelligence for incident root-cause analysis and forecasting. Develop LLM- and agent-based systems to improve operational efficiency, enhance LLM pipelines, and create robust data workflows for large-scale model development. Serve as a resident AI frameworks authority, applying automation, monitoring, and containerized (Kubernetes) cloud (AWS) implementations to sustain long-term reliability and product evolution.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 days ago

Senior AI Tools Engineer, SRE Operations - GeForce NOW

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and deploy AI-powered tools that transform GeForce NOW production signals, metrics, and logs into actionable intelligence for incident root-cause analysis and forecasting. Develop LLM- and agent-based systems to improve operational efficiency, enhance LLM pipelines, and create robust data workflows for large-scale model development. Serve as a resident AI frameworks authority, applying automation, monitoring, and containerized (Kubernetes) cloud (AWS) implementations to sustain long-term reliability and product evolution.
Location: Santa Clara, California
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and deploy sophisticated AI-powered tools and products for operating and optimizing the production GeForce NOW service.
  • •Transform production data streams (signals, metrics, logs) into actionable intelligence to automate root cause analysis and predict trends.
  • •Develop new LLM- and agent-based systems to improve operational efficiency.
  • •Establish and maintain data management practices, including workflows for large-scale data sources used in model development.
  • •Lead and enhance LLM-based pipelines and recommend AI frameworks, platforms, and architectural approaches for long-term technical sustainability.

Pay and Benefits

Salary: USD 144,000 - 230,000 annually
Equity and Bonus:Equity

Key Requirements

  • •B.S. in Computer Science, Statistics, or Engineering (or equivalent experience) and 5+ years of experience.
  • •Strong proficiency in Python; familiarity with Go or other systems languages is a plus.
  • •Practical experience building, optimizing, and deploying AI tools for production workflows.
  • •Strong knowledge of AI developments, including how LLM-based platforms are built and optimized.
  • •Hands-on experience with Kubernetes and cloud environments (AWS cloud).
Experience:5+ yearsSite reliability engineeringAILLMCloud
Education:Bachelor's in Computer Science, Statistics, or Engineering
Tech Stack:PythonGoKubernetesAWSGrafana

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor