Senior Site Reliability Engineer in Test, SDET

NVIDIA
Shanghai
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: mastersSkills: ["Problem-solving","Clear written communication","Clear verbal communication","Collaboration","Self-motivated"]

Design, build, and operate highly reliable test environments and automation infrastructure that enable validation of enterprise offerings. Own end-to-end CI/CD pipelines and GitOps-based CD using GitLab CI, GitHub Actions, and ArgoCD. Manage SBOM generation and policy enforcement, provision Kubernetes-based ephemeral environments, define reliability practices (SLIs/SLOs, error budgets), and harden pipelines through root-cause analysis and remediation. Apply AI/ML techniques to reduce flakiness and improve QA velocity.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Site Reliability Engineer in Test, SDET

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design, build, and operate highly reliable test environments and automation infrastructure that enable validation of enterprise offerings. Own end-to-end CI/CD pipelines and GitOps-based CD using GitLab CI, GitHub Actions, and ArgoCD. Manage SBOM generation and policy enforcement, provision Kubernetes-based ephemeral environments, define reliability practices (SLIs/SLOs, error budgets), and harden pipelines through root-cause analysis and remediation. Apply AI/ML techniques to reduce flakiness and improve QA velocity.
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, operate, and continuously improve reliable, scalable test environments and automation infrastructure for enterprise validation.
  • •Own end-to-end CI/CD pipelines using GitLab CI, GitHub Actions, and ArgoCD (GitOps), focusing on reliability, performance, and progressive delivery of test workloads.
  • •Manage SBOMs end-to-end (generation, continuous monitoring, vulnerability correlation, policy enforcement) and integrate them into CI/CD and release gates.
  • •Provision, scale, observe, and lifecycle-manage ephemeral and long-lived test environments with strong emphasis on isolation, reproducibility, and rapid recovery.
  • •Define and drive reliability practices for test systems (SLIs/SLOs, error budgets, toil reduction, chaos/resilience testing, automated remediation) and collaborate on triage and root-cause analysis.

Key Requirements

  • •MS or PhD in Computer Science (or equivalent) with 8+ years in Site Reliability Engineering, test environment management, CI/CD platform engineering, or software testing infrastructure.
  • •Strong Linux proficiency plus shell scripting and Python (or equivalent automation languages).
  • •Hands-on CI/CD system design and operation with GitLab CI and/or GitHub Actions, and practical ArgoCD (or equivalent GitOps) for CD of applications and infrastructure.
  • •Containerization and orchestration experience with Docker and Kubernetes (plus virtualization technologies).
  • •Deep SRE knowledge (SLIs/SLOs, error budgets, incident response, postmortems, toil elimination) and experience building test environments focused on reliability, isolation, and rapid turnaround.
Experience:8+ yearsSite reliability engineeringCI/CDTest environment managementTest automationDistributed systemsKubernetesGitOpsQuality assuranceAI/ML
Education:Master's
Skills:Problem-solvingClear written communicationClear verbal communicationCollaborationSelf-motivated
Tech Stack:LinuxShell scriptingPythonGitLab CIGitHub ActionsArgoCDGitOpsSBOMSBOM toolingOPAGatekeeperKyvernoDockerKubernetesCI/CDSLIsSLOsError budgetsGitLabGitHub

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor