DevOps Engineer - Équipe de Site Reliability Engineering de BTP (Montreal, Québec, CA, H3B 0B3)

SAP
Montreal
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsSkills: ["Teamwork","Emergency responsiveness","Problem solving","Communication"]

Operate and improve critical SAP and customer services with a strong focus on Site Reliability Engineering. Proactively monitor service behavior, investigate production incidents, and perform deep RCA to prevent recurrence. Build observability and recovery tooling using modern open-source and SAP technologies, partner closely with development teams on post-mortem improvements, and maintain technical documentation while following SRE best practices and on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SAP
SAP
2 hours ago

DevOps Engineer - Équipe de Site Reliability Engineering de BTP (Montreal, Québec, CA, H3B 0B3)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Operate and improve critical SAP and customer services with a strong focus on Site Reliability Engineering. Proactively monitor service behavior, investigate production incidents, and perform deep RCA to prevent recurrence. Build observability and recovery tooling using modern open-source and SAP technologies, partner closely with development teams on post-mortem improvements, and maintain technical documentation while following SRE best practices and on-call rotation.
Location: Montreal
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Serve as a technical expert during production incidents, investigating and resolving issues at a deep technical level.
  • •Conduct root-cause analyses (RCA) and drive follow-up improvements to prevent recurrence.
  • •Investigate complex problems using event and log analyses while respecting service level agreements (SLA).
  • •Design software solutions to improve reliability and stability, and enhance infrastructure/platform monitoring using key system metrics.
  • •Collaborate with development teams to implement improvements from post-mortems and maintain SRE documentation and best practices, including on-call rotation.
Travel: Low travel

Key Requirements

  • •2+ years of experience in a Site Reliability Engineering (SRE) role.
  • •Strong hands-on experience with Kubernetes and container technologies.
  • •Working knowledge of modern cloud architectures (AWS, Azure, or GCP is a plus).
  • •Proficiency with Linux/Unix systems and scripting, plus CI/CD experience (ArgoCD, Concourse, GitHub Actions are a plus).
  • •Experience with SRE tooling and incident investigation, including monitoring/logging/alerting (e.g., Grafana, Prometheus, Kibana, Loki, Splunk On-Call, Dynatrace) and security best practices for cloud app development.
Experience:2+ years
Skills:TeamworkEmergency responsivenessProblem solvingCommunication
Certifications:CKACKADCKS
Languages:EnglishFrench
Tech Stack:KubernetesLinuxUnixAWSAzureGCPArgoCDConcourseGitHub ActionsClaude Code CLIGitHub CopilotPythonGoBashGrafanaPrometheusKibanaLokiSplunk On-CallDynatrace

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

SAP
Global enterprise software company best known for ERP systems and business applications covering finance, supply chain, procurement, HR, analytics, and customer management for large organizations.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Walldorf, Germany
Founded: 1972
Glassdoor
Glassdoor: 4.1
WebsiteLinkedInGlassdoor