Senior Production Engineer, Core PE

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull time172,000 - 209,000Function: Manufacturing & Production OperationsExperience: 5+ yearsSkills: ["Communication","Collaboration","Problem-solving","Mentorship","Leadership"]

Join Crusoe as a Production Engineer focused on Operational Excellence to ensure reliability, scalability, and high performance of a GPU-based AI cloud. You will define availability metrics (SLIs/SLOs), drive observability and automation, lead incident response and RCAs, and collaborate across compute, networking, and platform teams to build self-healing infrastructure for massive AI workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
3 months ago

Senior Production Engineer, Core PE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join Crusoe as a Production Engineer focused on Operational Excellence to ensure reliability, scalability, and high performance of a GPU-based AI cloud. You will define availability metrics (SLIs/SLOs), drive observability and automation, lead incident response and RCAs, and collaborate across compute, networking, and platform teams to build self-healing infrastructure for massive AI workloads.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and evolve availability metrics for Crusoe’s cloud platform, including establishing, measuring, and improving SLIs and SLOs
  • •Participate in production incident response, diagnosing and resolving service disruptions while contributing to post-incident reviews and root cause analysis
  • •Build, operate, and improve observability across Crusoe’s infrastructure using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry
  • •Identify reliability risks, performance bottlenecks, and early indicators of potential production issues across distributed systems
  • •Develop automation and tooling that reduces operational toil, improves recovery times, and enables self-healing infrastructure

Pay and Benefits

Salary: 172,000 - 209,000
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kPaid ParentalCell PhoneEquity

Key Requirements

  • •5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
  • •Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems
  • •Strong knowledge of Linux/Unix systems, including debugging complex issues across kernel and user space
  • •Familiarity with infrastructure-as-code tools such as Terraform or Ansible
  • •Scripting or programming experience with Go, Python, C, or C++
Experience:5+ yearsGPUHPCDistributed systemsAI infrastructure
Skills:CommunicationCollaborationProblem-solvingMentorshipLeadership
Tech Stack:PrometheusGrafanaOpenTelemetryTerraformAnsibleKubernetesGoPythonC++CLinuxUnixAWSGCP

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor