DevOps Engineer - BTP Site Reliability Engineering team (Montreal, Quebec, CA, H3B 0B3)

SAP
Montreal
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsSkills: ["Team player","Self-motivated","Driven","Communication","Root cause analysis"]

Join an established Site Reliability Engineering team for the SAP Business AI Platform, delivering 24x7 deep technical coverage. You’ll monitor and troubleshoot critical cloud services, lead incident response and root-cause analysis, and build reliability improvements using SRE principles. Work with development teams on postmortems and product improvements, enhance monitoring with the four Golden Signals, and support on-call rotation for major incidents.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SAP
SAP
2 hours ago

DevOps Engineer - BTP Site Reliability Engineering team (Montreal, Quebec, CA, H3B 0B3)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Join an established Site Reliability Engineering team for the SAP Business AI Platform, delivering 24x7 deep technical coverage. You’ll monitor and troubleshoot critical cloud services, lead incident response and root-cause analysis, and build reliability improvements using SRE principles. Work with development teams on postmortems and product improvements, enhance monitoring with the four Golden Signals, and support on-call rotation for major incidents.
Location: Montreal
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Respond as technical expert during live site incidents, investigating and solving complex outages and service-impacting issues.
  • •Drive root cause analysis and implement follow-up improvements to prevent recurrence.
  • •Perform deep troubleshooting and log analysis to resolve issues in line with internal and external SLAs.
  • •Build and enhance software and tooling for service reliability, monitoring, and recovery using SRE principles.
  • •Collaborate with development teams on postmortems, product improvements, and maintain technical documentation while participating in on-call rotation.
Travel: Low travel

Key Requirements

  • •2+ years of experience in SRE.
  • •Experience with Kubernetes and container technologies.
  • •Hands-on Unix/Linux experience and scripting skills.
  • •CI/CD experience (ArgoCD, Concourse, or GitHub Actions) and strong automation mindset.
  • •Fluency in English and ability to work efficiently during emergencies in a global team setup.
Experience:2+ years
Skills:Team playerSelf-motivatedDrivenCommunicationRoot cause analysis
Certifications:CKACKADCKS
Languages:English
Tech Stack:KubernetesAWSAzureGCPUnix/LinuxArgoCDConcourseGitHub ActionsClaude Code CLIGitHub CopilotPythonGoBashGrafanaPrometheusKibanaLokiSplunk On-CallDynatrace

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

SAP
Global enterprise software company best known for ERP systems and business applications covering finance, supply chain, procurement, HR, analytics, and customer management for large organizations.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Walldorf, Germany
Founded: 1972
Glassdoor
Glassdoor: 4.1
WebsiteLinkedInGlassdoor