Principal Engineer, Cloud Site Reliability Engineering

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 272,000 - 431,250 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 15+ yearsEducation: bachelorsSkills: ["Collaboration","Interpersonal skills","Guiding and influencing","Alignment","Problem-solving"]

Design and architect SRE solutions for GPU Private Cloud used by thousands of NVIDIA engineers for interactive development, centralized CI/CD, and QA testing. You’ll evaluate and optimize software development workflows, build end-to-end CI/CD pipelines using open-source and proprietary tooling, onboard internal teams, and improve AI development speed and cost. Lead technical projects, resolve system issues, and implement critical metrics, dashboards, and analytics.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Principal Engineer, Cloud Site Reliability Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design and architect SRE solutions for GPU Private Cloud used by thousands of NVIDIA engineers for interactive development, centralized CI/CD, and QA testing. You’ll evaluate and optimize software development workflows, build end-to-end CI/CD pipelines using open-source and proprietary tooling, onboard internal teams, and improve AI development speed and cost. Lead technical projects, resolve system issues, and implement critical metrics, dashboards, and analytics.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Serve as an SRE Architect for the GPU Private Cloud team supporting thousands of NVIDIA users.
  • •Architect, implement, and support end-to-end CI/CD systems using open-source and NVIDIA proprietary software.
  • •Evaluate and develop solutions to optimize critical software development workflows across NVIDIA organizations.
  • •Onboard internal development teams to private cloud infrastructure by discovering use cases and available solutions.
  • •Identify performance bottlenecks and optimize speed and cost efficiency for AI development and testing, including implementing metrics and dashboards.

Pay and Benefits

Salary: USD 272,000 - 431,250 annually
Equity and Bonus:Equity

Key Requirements

  • •BS or MS in Electrical Engineering, Computer Science, or related field (or equivalent experience).
  • •15+ years of systems software development, including at least 1 year developing or exploring AI.
  • •Experience maintaining cloud infrastructure and highly available production environments.
  • •Strong programming skills in Java, Python, and Shell scripting, with understanding of distributed systems and REST APIs.
  • •Experience with SQL/NoSQL databases and production container/VM environments, including Docker; Kubernetes and related cloud tools are required.
Experience:15+ yearsAIDistributed systemsCloud infrastructure
Education:Bachelor's
Skills:CollaborationInterpersonal skillsGuiding and influencingAlignmentProblem-solving
Tech Stack:JavaPythonShell-scriptREST APIsSQLNoSQLMySQLCassandraMongoDBElasticsearchDockerVirtual MachinesOpenStackKubernetesChefPuppetHadoopCephSwiftStackLXC

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor