Staff Production Engineer, Core PE

Crusoe
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Strong communication","Problem-solving","Team collaboration","Automation","Incident management"]

Site Reliability Engineer focused on Operational Excellence at Crusoe, you will ensure stability, resilience, and performance of the GPU cloud; partner with SREs and platform teams to reduce toil, improve incident response, and drive reliability across a large-scale distributed platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
9 months ago

Staff Production Engineer, Core PE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Site Reliability Engineer focused on Operational Excellence at Crusoe, you will ensure stability, resilience, and performance of the GPU cloud; partner with SREs and platform teams to reduce toil, improve incident response, and drive reliability across a large-scale distributed platform.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and refine availability metrics for Crusoe’s cloud infrastructure, including establishing, tracking, and improving SLIs and SLOs.
  • •Assist in incident response by identifying, diagnosing, and resolving service disruptions, and support post-incident processes through RCA documentation and post-incident reviews.
  • •Build, operate, and monitor infrastructure health using Crusoe’s observability stack (Prometheus, Grafana, Alertmanager, OpenTelemetry).
  • •Identify and communicate reliability risks, performance bottlenecks, and early indicators of potential incidents that could impact service availability.
  • •Develop automation and tooling to reduce operational toil, minimize manual intervention, and enhance service recovery and self-healing capabilities.

Pay and Benefits

Perks:Health Insurance401kPaid LeaveParental LeaveLife InsuranceDisability

Key Requirements

  • •5+ years of experience in cloud operations, SRE, or related roles.
  • •Understanding of cloud platforms and infrastructure fundamentals (Kubernetes, AWS/GCP, virtualization, distributed systems)
  • •Familiarity with incident management practices and operational frameworks (SRE/ITIL/etc.)
  • •Experience with monitoring and alerting tools (Prometheus, Grafana)
  • •Familiarity with infrastructure-as-code and configuration management tools such as Terraform and Ansible
Experience:5+ yearsCloudSREDistributed systems
Skills:Strong communicationProblem-solvingTeam collaborationAutomationIncident management
Tech Stack:PrometheusGrafanaOpenTelemetryKubernetesAWSGCPTerraformAnsibleGoPython

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor