Site Reliability Engineer – Manchester

BAE Systems
United Kingdom
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Troubleshooting","Problem-solving","Automation mindset","Collaboration","Incident response"]

Support and maintain essential mission applications by applying software and systems engineering to automate operations and improve availability, performance, and stability. Instrument services lacking monitoring, diagnose outages across the stack, and advise product teams on reliable system design. You’ll work in an Agile/Scrum environment using tools such as Jira, deployment automation (Chef/Puppet), monitoring (ELK), and container/microservices patterns with Docker while contributing to the wider DevOps/SRE community.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
BAE Systems
BAE Systems
2 months ago

Site Reliability Engineer – Manchester

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Support and maintain essential mission applications by applying software and systems engineering to automate operations and improve availability, performance, and stability. Instrument services lacking monitoring, diagnose outages across the stack, and advise product teams on reliable system design. You’ll work in an Agile/Scrum environment using tools such as Jira, deployment automation (Chef/Puppet), monitoring (ELK), and container/microservices patterns with Docker while contributing to the wider DevOps/SRE community.
Location: United Kingdom
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Support and maintain essential services for core mission applications, improving availability, performance, and stability.
  • •Automate operational work to reduce manual operations such as incident tickets and on-call, keeping support time to no more than half of the team’s time.
  • •Instrument applications lacking monitoring and use outputs to demonstrate daily improvement to the team’s reliability impact.
  • •Partner with product teams to advise on best practices for designing and building scalable, resilient systems.
  • •Diagnose and troubleshoot application issues causing service outages and participate in the wider DevOps/SRE community.

Key Requirements

  • •Use software development in Java and web technologies (e.g., JavaScript, HTML) to support production services.
  • •Have hands-on Linux and Windows command-line skills (e.g., Bash, PowerShell) and strong troubleshooting across the stack.
  • •Work with cloud infrastructure such as AWS, Azure, or OpenStack and deployment tools like Chef and Puppet.
  • •Use monitoring for large systems using technologies such as ELK and instrument applications to improve observability.
  • •Experience with container management and micro-services architectures (e.g., Docker), plus familiarity with automation and agile tools such as Jira.
Skills:TroubleshootingProblem-solvingAutomation mindsetCollaborationIncident response
Tech Stack:JavaJavaScriptHTMLElasticMongoLinuxWindowsBashPowerShellAWSAzureOpenStackChefPuppetELKJiraDockerMicro-servicesOpen SourceSelenium

Eligibility

Security Clearance:Baseline Personnel Security Standard

Company Brief

BAE Systems
Multinational defence, security and aerospace company designing, manufacturing and supporting military and civilian systems including naval ships, submarines, aircraft, cyber solutions and electronic systems for governments and commercial customers.
Industry: Defense Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: London, United Kingdom
Founded: 1999
WebsiteLinkedIn