Site Reliability Engineer

Oxio
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Continuous improvement","Incident management","Scripting","Observability","Troubleshooting"]

Design and implement cloud platform capabilities that support backend services, including automated deployments, scaling, recovery, monitoring, and uptime-focused operations. Own production reliability through incident response, on-call rotation, and blameless postmortems. Provide engineering teams with tools to operate the service they build, enabling continuous improvement across infrastructure and telecom platform operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Oxio
Oxio
2 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live
Reposted: similar role first listed 3 months ago

Job Summary

Design and implement cloud platform capabilities that support backend services, including automated deployments, scaling, recovery, monitoring, and uptime-focused operations. Own production reliability through incident response, on-call rotation, and blameless postmortems. Provide engineering teams with tools to operate the service they build, enabling continuous improvement across infrastructure and telecom platform operations.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design and implement cloud platform capabilities to support backend services.
  • •Automate technical operations including deployments, scaling, and recovery.
  • •Monitor and maintain mission-critical production infrastructure to ensure maximum uptime.
  • •Participate in an on-call rotation and drive continuous improvement through blameless postmortems.
  • •Enable Engineering/Telecom/Data Engineering teams by providing tools to operate the service.

Key Requirements

  • •Understanding of Linux/Unix systems and internals such as process management, filesystems, memory management, and networking.
  • •Proficiency in at least one programming language (Python, Go, or Ruby) plus strong scripting skills (Bash, Perl).
  • •Experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible.
  • •Familiarity with containerization and orchestration, including Docker and Kubernetes.
  • •Experience with monitoring/observability and incident practices, including alerts, log analysis, dashboards, runbooks, and postmortems (plus on-call experience).
Skills:Continuous improvementIncident managementScriptingObservabilityTroubleshooting
Tech Stack:LinuxUnixPythonGoRubyBashPerlTerraformCloudFormationAnsibleDockerKubernetesPrometheusGrafanaDatadogJenkinsGitLab CICircleCIAWSGoogle Cloud

Company Brief

Oxio
OXIO provides a cloud-native Telecom-as-a-Service (TaaS) platform that enables brands and enterprises to embed and operate mobile connectivity, MVNO services, and data-driven telecom products via APIs and a programmable network.
Industry: Telecommunications
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series B
Headquarters: New York, United States
Founded: 2018
Glassdoor
Glassdoor: 3.2
WebsiteLinkedInGlassdoor