Site Reliability Engineer

Crusoe
Dublin
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 1-3 yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Collaboration"]

Join Crusoe’s SRE team to ensure high reliability and performance of the AI platform. You’ll automate infrastructure, build internal tooling, monitor systems, respond to incidents, and collaborate with software engineers to implement resilient, scalable services across distributed environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
7 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join Crusoe’s SRE team to ensure high reliability and performance of the AI platform. You’ll automate infrastructure, build internal tooling, monitor systems, respond to incidents, and collaborate with software engineers to implement resilient, scalable services across distributed environments.
Location: Dublin
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Entry level

Key Responsibilities

  • •Automate routine processes and build Crusoe’s internal infrastructure platform to enable software teams to run their services without needing to understand the OS, hardware, or network.
  • •Collaborate in daily stand-ups and with software engineers to review deployment readiness and plan data center deployments or retrofits.
  • •Review overnight alerts and system performance metrics to ensure reliability and analyze logs to improve monitoring capabilities.
  • •Engage in incident response drills, post-mortems, and root cause analyses to prevent recurrence and automate common errors.
  • •Maintain high SLIs/SLOs and document work for knowledge sharing and future planning.

Pay and Benefits

Perks:PensionHealth InsuranceDentalIncome ProtectionLife Assurance

Key Requirements

  • •1-3 years of professional SRE experience
  • •Exposure to server-class hardware and provisioning
  • •Understanding of distributed system architecture and reliability patterns
  • •Proficiency in at least one programming language (Python, Go, or similar)
  • •Familiarity with Docker, Kubernetes, Ansible, CloudFormation, Terraform and CI/CD tools (Jenkins, GitLab, CircleCI, GitHub Actions)
Experience:1-3 yearsCloud infrastructureDistributed systemsSREDevOps
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaboration
Tech Stack:PythonGoDockerKubernetesAnsibleCloudFormationTerraformJenkinsGitLabCircleCIGitHub ActionsUnixLinuxTCP/IP

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor