Site Reliability Engineer - SAP Business AI Platform (BAIP)

SAP
Sofia
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Troubleshooting","Problem-solving","Curiosity","Teamwork"]

Own operational excellence for business-critical SAP Business AI Platform services, applying SRE practices such as SLO/SLI management, incident response, RCA follow-ups, and chaos engineering. Troubleshoot production issues using distributed tracing and Kubernetes debugging, improve alerting quality to reduce noise, and build automation and tooling to eliminate toil. Collaborate with development teams, contribute to knowledge sharing, and leverage SAP’s AI-first tooling for SRE workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SAP
SAP
2 hours ago

Site Reliability Engineer - SAP Business AI Platform (BAIP)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own operational excellence for business-critical SAP Business AI Platform services, applying SRE practices such as SLO/SLI management, incident response, RCA follow-ups, and chaos engineering. Troubleshoot production issues using distributed tracing and Kubernetes debugging, improve alerting quality to reduce noise, and build automation and tooling to eliminate toil. Collaborate with development teams, contribute to knowledge sharing, and leverage SAP’s AI-first tooling for SRE workflows.
Location: Sofia
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Participate in hotline and on-call rotation, responding to incidents and SLO violations.
  • •Perform deep technical analysis of production issues using log analysis, distributed tracing, and Kubernetes debugging.
  • •Own reliability engineering activities: monitor SLO/SLIs and 4 Golden Signals, and drive improvements to alerting quality.
  • •Contribute to automation and tooling (scripts, Recommended Actions, runbooks, and internal tools) to reduce toil.
  • •Partner with development teams on reliability improvements, participate in chaos/fail drills, and support onboarding new services into SRE scope.
Travel: Low travel

Key Requirements

  • •Solid Linux/Unix foundation and comfort operating in a terminal with strong system administration fundamentals.
  • •Networking fundamentals including TCP/IP, DNS, HTTP/S, load balancing, and basic troubleshooting.
  • •Curiosity about distributed systems failure modes and how to make services more resilient.
  • •Ability to communicate clearly in incidents, writing, and cross-team discussions; fluent English.
  • •Team-first mindset with on-call responsibility and strong troubleshooting/problem-solving skills.
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:CommunicationTroubleshootingProblem-solvingCuriosityTeamwork
Languages:English
Tech Stack:LinuxUnixKubernetesK8sHelmIstioArgoCDCloud FoundryAWSGCPAzureTerraformConcourseJenkinsVaultGardenerKymaDynatraceELK StackElasticsearch

Company Brief

SAP
Global enterprise software company best known for ERP systems and business applications covering finance, supply chain, procurement, HR, analytics, and customer management for large organizations.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Walldorf, Germany
Founded: 1972
Glassdoor
Glassdoor: 4.1
WebsiteLinkedInGlassdoor