Site Reliability Engineer Engineer

Modus-create
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Troubleshooting","Mentoring","Stakeholder management"]

Build reliable, observable systems for client teams by improving uptime, performance, and incident prevention. Partner with developers, platform engineers, and operations to reduce toil and shift toward proactive reliability engineering. Own 24/7 P0 on-call with a 10-minute acknowledgment SLA, lead incident investigations and postmortems, and mentor SRE engineers as the team grows. Work across AWS infrastructure and Python backend services, implementing observability, SLOs, and infrastructure-as-code.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modus-create
Modus-create
19 hours ago

Site Reliability Engineer Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build reliable, observable systems for client teams by improving uptime, performance, and incident prevention. Partner with developers, platform engineers, and operations to reduce toil and shift toward proactive reliability engineering. Own 24/7 P0 on-call with a 10-minute acknowledgment SLA, lead incident investigations and postmortems, and mentor SRE engineers as the team grows. Work across AWS infrastructure and Python backend services, implementing observability, SLOs, and infrastructure-as-code.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design observability systems (metrics, logging, tracing) and define SLOs, error budgets, and monitoring strategies
  • •Own 24/7 P0 on-call rotation with a 10-minute acknowledgment SLA and validate/escalate incident reports
  • •Establish reliability standards and SLA targets for backend services
  • •Mentor and onboard additional SRE engineers as the team scales from 4 to 8-12
  • •Participate in incident response, troubleshooting, root-cause analysis, and postmortems

Key Requirements

  • •6+ years hands-on experience in SRE, DevOps, or Platform Engineering at scale
  • •Deep AWS expertise across services such as ALB, ECS/Fargate, RDS/Aurora, Lambda, and IAM
  • •Production Python backend engineering experience, including debugging and optimizing services in containers and Lambdas
  • •Experience building or bootstrapping SRE programs and operating incident response/on-call with SLOs and postmortems
  • •Deep experience designing and implementing observability (metrics, logging, tracing, alerting)
Skills:CommunicationTroubleshootingMentoringStakeholder management
Languages:English
Tech Stack:AWSALBECS/FargateAuroraRDSLambdaIAMPythonKubernetesDockerTerraformCloudFormationPulumiCI/CDLinux/UnixBashGoGitJiraAtlassian

Company Brief

Modus-create
Digital transformation consultancy delivering product design, cloud-native engineering, DevOps, and organizational change services to help enterprises build and scale modern digital products and platforms.
Industry: Consulting
Company Size: Medium (51 to 250 employees)
Growth: Established Company
Funding: Bootstrapped
Headquarters: Raleigh, United States
Founded: 2007
WebsiteLinkedIn