Engenheiro(a) Sênior de Confiabilidade de Sites (SRE) – Storage

Dell
Rio Grande do Sul
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Analytical thinking","Problem-solving","Troubleshooting","Collaboration","Incident analysis"]

Design, develop, and maintain automated storage-service lifecycle workflows, from provisioning and capacity management to replication integrity monitoring, firmware management, and decommissioning—run via pipelines without direct production array interaction. Operate reliability practices using SLIs, SLOs, and error budgets, improve observability with monitoring and telemetry tools, and perform structured incident analyses. Partner across Storage Engineering, Platform Automation, and Observability to integrate automation into the corporate AI/agentic ecosystem and drive continuous improvement.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Dell
Dell
1 month ago

Engenheiro(a) Sênior de Confiabilidade de Sites (SRE) – Storage

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Design, develop, and maintain automated storage-service lifecycle workflows, from provisioning and capacity management to replication integrity monitoring, firmware management, and decommissioning—run via pipelines without direct production array interaction. Operate reliability practices using SLIs, SLOs, and error budgets, improve observability with monitoring and telemetry tools, and perform structured incident analyses. Partner across Storage Engineering, Platform Automation, and Observability to integrate automation into the corporate AI/agentic ecosystem and drive continuous improvement.
Location: Rio Grande do Sul
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, develop, and maintain automations for the full lifecycle of storage services, including volume/LUN provisioning, snapshot and replication management, capacity alerts, firmware management, and automated remediation.
  • •Define and operate reliability practices for storage services using SLIs, SLOs, and error budgets, and enhance observability with monitoring/telemetry tools; conduct structured post-incident root-cause analyses and automate recurring failures.
  • •Collaborate with Storage Engineering, Platform Automation, and Observability teams to set operational requirements, monitoring baselines, supported I/O profiles, replication tolerances, capacity limits, readiness criteria, and integrate automation with the corporate agentic AI platform.
  • •Drive continuous improvement by measuring automation coverage and autonomous resolution rates, maintaining structured knowledge bases, documenting resolved incidents, and identifying/prioritizing automation opportunities.
  • •Participate in Storage tower on-call escalation, using consolidated diagnostics (arrays status, replication state, capacity metrics, telemetry, change history, and recommended actions) to resolve critical incidents within service levels.

Key Requirements

  • •Complete higher education in Computer Science, IT, Engineering, or related fields, with experience in Storage Engineering, Storage Operations, SRE, or large-scale corporate infrastructure.
  • •Advanced or fluent English, with ability to collaborate effectively with global teams.
  • •At least 5 years of experience in Storage Engineering or Storage Operations in large enterprise environments.
  • •Advanced hands-on experience in at least two storage areas such as SAN/Block, NAS/File Storage, Object Storage (including listed vendor platforms and/or S3-compatible technologies).
  • •Proven experience building storage automation using Ansible, Python, REST APIs, Git, CI/CD pipelines, and GitOps practices (covering provisioning, replication, capacity management, monitoring, and lifecycle).
Experience:5+ yearsStorage engineeringSite reliability engineering (SRE)Enterprise storage operationsMultivendor storageS3-compatible platforms
Education:Bachelor's in Computer Science, Technology of Information, Engineering or related fields
Skills:Analytical thinkingProblem-solvingTroubleshootingCollaborationIncident analysis
Languages:English
Tech Stack:AnsiblePythonREST APIsGitCI/CDGitOpsServiceNow CMDBGitLab CIZabbixVROpsMoogsoftCriblSIEMData lakehouseSLIsSLOsError budgetsTelemetryLLMsPolicy-as-code

Company Brief

Dell
Designs, manufactures and sells computers, servers, storage, networking equipment, and related IT solutions and services for consumers, small businesses, and enterprises worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Round Rock, United States
Founded: 1984
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor