Site Reliability Engineer Engineer

Modus-create
Costa Rica
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsSkills: ["Incident response","Communication","Mentoring","Troubleshooting","Collaboration"]

Build reliable, observable systems that help client teams operate with confidence at scale. Work embedded with cross-functional teams to reduce toil, prevent incidents, and improve uptime and performance across AWS infrastructure and Python backend services. Own 24/7 P0 on-call, drive incident response and postmortems, define SLOs and error budgets, and implement reliability improvements. Mentor and onboard SRE engineers as the team grows, while advising on reliability strategy and operational excellence.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modus-create
Modus-create
15 hours ago

Site Reliability Engineer Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build reliable, observable systems that help client teams operate with confidence at scale. Work embedded with cross-functional teams to reduce toil, prevent incidents, and improve uptime and performance across AWS infrastructure and Python backend services. Own 24/7 P0 on-call, drive incident response and postmortems, define SLOs and error budgets, and implement reliability improvements. Mentor and onboard SRE engineers as the team grows, while advising on reliability strategy and operational excellence.
Location: Costa Rica
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design observability systems (metrics, logging, tracing) and define SLOs, error budgets, and monitoring strategies.
  • •Own the 24/7 P0 on-call rotation with a 10-minute acknowledgment SLA, validate/escalate AI-generated incident reports, and manage incident response.
  • •Lead incident investigations and postmortems, and drive reliability improvements, capacity planning, and performance optimization.
  • •Establish reliability standards and SLA targets for backend services, tuning for latency, throughput, and resource efficiency.
  • •Mentor and onboard additional SRE engineers as the team scales from 4 to 8-12 over the coming months.

Key Requirements

  • •6+ years hands-on experience in SRE, DevOps, or Platform Engineering at scale.
  • •Deep AWS expertise (ALB, ECS/Fargate, RDS/Aurora, Lambda, IAM) or strong proficiency in a cloud platform.
  • •Production Python backend engineering experience for debugging and optimizing services in containers and Lambdas.
  • •Experience building or bootstrapping SRE programs from scratch, including incident response, on-call rotations, postmortems, and SLOs.
  • •Hands-on observability experience (metrics, logging, tracing, alerting) and infrastructure-as-code (Terraform, CloudFormation, Pulumi, etc.).
Experience:6+ years
Skills:Incident responseCommunicationMentoringTroubleshootingCollaboration
Tech Stack:AWSALBECS/FargateAuroraRDSLambdaIAMPythonKubernetesDockerTerraformCloudFormationPulumiCI/CDLinux/UnixBashGoGitAtlassianJira

Company Brief

Modus-create
Digital transformation consultancy delivering product design, cloud-native engineering, DevOps, and organizational change services to help enterprises build and scale modern digital products and platforms.
Industry: Consulting
Company Size: Medium (51 to 250 employees)
Growth: Established Company
Funding: Bootstrapped
Headquarters: Raleigh, United States
Founded: 2007
WebsiteLinkedIn