Site Reliability Operations Engineer

Salesforce
Seattle
Workplace: OnsiteFull timeUSD 94,000 - 142,300 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5-8 yearsSkills: ["Incident management","Problem management","Communication","Troubleshooting","Automation"]

Own major incident response for internal Salesforce systems as the Incident Commander, coordinating technical teams to restore service quickly. Monitor and troubleshoot enterprise infrastructure, applications, and networking across platforms and vendors, improving reliability through runbooks, SOPs, automation, and emergency change coordination. Analyze incident data and KPIs to reduce impact duration, lead problem management and root-cause investigations, and participate in regional on-call coverage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Salesforce
Salesforce
2 hours ago

Site Reliability Operations Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own major incident response for internal Salesforce systems as the Incident Commander, coordinating technical teams to restore service quickly. Monitor and troubleshoot enterprise infrastructure, applications, and networking across platforms and vendors, improving reliability through runbooks, SOPs, automation, and emergency change coordination. Analyze incident data and KPIs to reduce impact duration, lead problem management and root-cause investigations, and participate in regional on-call coverage.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Respond to and manage major incidents, serving as Incident Commander to coordinate teams and drive rapid service restoration.
  • •Monitor and troubleshoot enterprise systems across infrastructure, applications, and networking components before they impact users.
  • •Improve incident response globally by creating/updating runbooks and SOPs and driving automation.
  • •Coordinate emergency changes and infrastructure updates with cross-functional teams to maintain business continuity.
  • •Lead problem management and post-incident activities, including investigating recurring incidents and conducting root cause analyses.

Pay and Benefits

Salary: USD 94,000 - 142,300 annually
Perks:Health InsuranceDentalVisionPaid ParentalLife InsuranceDisability Insurance401kEquity

Key Requirements

  • •5-8 years in IT operations, incident management, or site reliability work, preferably in a 24x7 high-availability environment.
  • •Ability to manage high-severity incidents under pressure, establishing impact and balancing technical and business needs.
  • •Strong verbal and written communication for explaining complex technical issues to technical and executive audiences.
  • •Technical troubleshooting across Windows and Linux servers, networking, cloud platforms, and virtualization technologies, using logs and monitoring tools.
  • •A related technical degree is required.
Experience:5-8 years
Skills:Incident managementProblem managementCommunicationTroubleshootingAutomation
Certifications:ITILAWSCCNAMCSARHCE
Tech Stack:AWSPythonBashPowerShellWindowsLinuxNetworkingCloud platformsVirtualizationSplunkGrafanaTableauPuppetChefITIL

Company Brief

Salesforce
Provides a leading cloud-based customer relationship management (CRM) platform with sales, service, marketing, analytics, and integration tools that empower businesses to manage customer relationships and digital transformation at scale.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 1999
Glassdoor
Glassdoor: 4.1
WebsiteLinkedInGlassdoor