Manager, Production Engineering

Crusoe
Tel Aviv
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Leadership","Mentorship","Problem-solving","Sense of urgency","Hands-on execution"]

Lead and build Crusoe’s Tel Aviv Production Engineering presence from scratch within the Cloud Availability Platform Engineering organization. This hybrid role combines people leadership with hands-on software and incident response to transition reliability from reactive firefighting to software-defined operations. Drive alert-reduction, predictive monitoring, and runbook automation (e.g., Temporal), while partnering cross-functionally across compute, storage, networking, and platform teams to enforce production readiness and change control.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 month ago

Manager, Production Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Lead and build Crusoe’s Tel Aviv Production Engineering presence from scratch within the Cloud Availability Platform Engineering organization. This hybrid role combines people leadership with hands-on software and incident response to transition reliability from reactive firefighting to software-defined operations. Drive alert-reduction, predictive monitoring, and runbook automation (e.g., Temporal), while partnering cross-functionally across compute, storage, networking, and platform teams to enforce production readiness and change control.
Location: Tel Aviv
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Recruit, mentor, and establish a high-performing Production Engineering footprint in Tel Aviv and set a culture of operational discipline.
  • •Hire, scale, and lead a local team while remaining hands-on in code and incident response.
  • •Partner with US and Dublin teams to run follow-the-sun global on-call and enforce blameless post-mortems targeting systemic failures.
  • •Improve reliability via alert-reduction, predictive monitoring for SEV1/SEV2 events, and automation of manual workflows (e.g., Temporal).
  • •Act as a Production Gatekeeper across cross-functional teams through Production Readiness Reviews and change control, including physical-to-digital auto-remediation.

Pay and Benefits

Perks:Pension

Key Requirements

  • •Minimum of 8+ years working in infrastructure, SRE, or production engineering environments.
  • •Minimum of 2+ years leading first-line engineering teams in a high-growth neocloud/hyperscaler or large-scale distributed environment.
  • •Hands-on coding proficiency in Go, Python, C++, or a comparable systems language to build automation.
  • •Expert command of Linux internals, container orchestration at scale, and root-cause analysis across physical-to-virtual boundaries.
  • •Proven experience running tiered on-call models with clear SLIs/SLOs and error budgets to reduce paging fatigue.
Experience:8+ yearsAI infrastructureNeocloudHyperscaler
Skills:LeadershipMentorshipProblem-solvingSense of urgencyHands-on execution
Tech Stack:GoPythonC++LinuxContainersRunbook automationTemporalInfiniBandRoCEv2BMCFirmware qualificationAttestationDCGMNVIDIAAMDGPU clusters

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor