Site Reliability Engineer Engineer

Modus-create
Brazil, Colombia, Mexico
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsSkills: ["Communication","Mentoring","Collaboration","Troubleshooting","Reliability engineering"]

Build reliable, observable systems that help client teams operate confidently at scale. Partner with developers and platform/operations teams to reduce toil, improve uptime, and prevent incidents through strong SRE practices. Own 24/7 P0 on-call with fast acknowledgments, lead incident response, root-cause analysis, and postmortems, and design observability with SLOs and error budgets. Mentor and onboard additional SREs as the team grows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modus-create
Modus-create
22 hours ago

Site Reliability Engineer Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build reliable, observable systems that help client teams operate confidently at scale. Partner with developers and platform/operations teams to reduce toil, improve uptime, and prevent incidents through strong SRE practices. Own 24/7 P0 on-call with fast acknowledgments, lead incident response, root-cause analysis, and postmortems, and design observability with SLOs and error budgets. Mentor and onboard additional SREs as the team grows.
Location: Brazil, Colombia, Mexico
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design observability systems (metrics, logging, tracing) and define SLOs, error budgets, and monitoring strategies.
  • •Own a 24/7 P0 on-call rotation with a 10-minute acknowledgment SLA, validate and escalate incident reports.
  • •Establish reliability standards and SLA targets for backend services.
  • •Lead incident response, troubleshooting, root-cause analysis, and postmortems.
  • •Implement reliability improvements, capacity planning, performance optimization, and self-healing/toil-reduction automation.

Key Requirements

  • •6+ years hands-on experience in SRE, DevOps, or Platform Engineering at scale.
  • •Strong AWS proficiency (ALB, ECS/Fargate, RDS/Aurora, Lambda, IAM).
  • •Production Python backend engineering experience in containers and Lambdas.
  • •Experience building or bootstrapping SRE programs, including incident response, on-call rotations, postmortems, and SLOs.
  • •Deep observability experience (metrics, logging, tracing, alerting) and infrastructure-as-code (Terraform, CloudFormation, Pulumi, etc.).
Experience:6+ years
Skills:CommunicationMentoringCollaborationTroubleshootingReliability engineering
Languages:English
Tech Stack:AWSALBECSFargateAuroraRDSLambdaIAMPythonKubernetesDockerTerraformCloudFormationPulumiCI/CDCI/CD pipelinesGitBashGoLinux/Unix

Company Brief

Modus-create
Digital transformation consultancy delivering product design, cloud-native engineering, DevOps, and organizational change services to help enterprises build and scale modern digital products and platforms.
Industry: Consulting
Company Size: Medium (51 to 250 employees)
Growth: Established Company
Funding: Bootstrapped
Headquarters: Raleigh, United States
Founded: 2007
WebsiteLinkedIn