Senior Site Reliability Engineer

Procter & Gamble
Philippines
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Problem-solving","Troubleshooting","Communication","Cross-functional collaboration","Leadership"]

Lead the SRE Incident Response team for Warehousing IT Operations, driving swift resolution of critical incidents while improving reliability, scalability, and performance. Own monitoring, alerting, and automated incident response; manage SLOs/SLIs and incident reporting to stakeholders. Collaborate with cross-functional engineering, DevOps, and site customers to design resilient systems, conduct root-cause analysis, and implement preventive measures. Provide mentorship, schedule coverage, and continuous upskilling within the team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Procter & Gamble
Procter & Gamble
3 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Lead the SRE Incident Response team for Warehousing IT Operations, driving swift resolution of critical incidents while improving reliability, scalability, and performance. Own monitoring, alerting, and automated incident response; manage SLOs/SLIs and incident reporting to stakeholders. Collaborate with cross-functional engineering, DevOps, and site customers to design resilient systems, conduct root-cause analysis, and implement preventive measures. Provide mentorship, schedule coverage, and continuous upskilling within the team.
Location: Philippines
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Lead incident response for critical system issues, minimizing downtime and user impact.
  • •Run incident management processes including clear communication, coordination, and documentation; perform root cause analysis and preventive improvements.
  • •Ensure high system availability via robust monitoring, alerting, and automated incident response systems.
  • •Implement comprehensive monitoring/observability solutions and manage SLOs/SLIs to meet service level objectives.
  • •Lead and mentor the SRE team, manage time/schedule with rotating on-call coverage, and report incident and performance metrics to stakeholders.

Pay and Benefits

Perks:Health InsuranceGym MembershipEmployee AssistanceRemote WorkAnnual Bonus

Key Requirements

  • •7+ years of industry experience in software engineering, software development, SRE, DevOps, or technical consulting.
  • •Knowledge of system administration on Linux/Unix, cloud platforms (AWS, Azure, or GCP), and SAP.
  • •Experience with configuration management and infrastructure-as-code frameworks such as Terraform.
  • •Proficiency in at least one programming language (Python or C#) and scripting for automation.
  • •Familiarity with incident response, root cause analysis, monitoring/observability tools (Prometheus, Grafana), and secure systems.
Experience:7+ yearsSoftware engineeringSREDevOpsTechnical consulting
Skills:Problem-solvingTroubleshootingCommunicationCross-functional collaborationLeadership
Tech Stack:Linux/UnixAWSAzureGCPSAPTerraformPythonC#NetworkingDNSLoad balancingDockerKubernetesSQLPrometheusGrafana

Company Brief

Procter & Gamble
Procter & Gamble is a global consumer goods company that develops, manufactures, and markets branded household, personal care, and health products across categories like beauty, grooming, health care, fabric & home care, and baby & family care.
Industry: Consumer Goods
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cincinnati, United States
Founded: 1837
WebsiteLinkedIn