Senior Production Engineer, Managed Cloud

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 170,000 - 205,000 annuallyFunction: Manufacturing & Production OperationsSkills: ["Collaboration","Communication","Problem-solving","Urgency","Reliability mindset"]

Build and operate reliable managed AI services for large-scale LLM workloads on an AI-optimized cloud platform. Define and improve SLIs/SLOs, automate observability with telemetry and performance tuning, and investigate reliability issues using logs and profiling. Partner with AI, platform, and infrastructure teams to optimize training and inference clusters, and help design next-generation distributed systems for AI-first environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 month ago

Senior Production Engineer, Managed Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and operate reliable managed AI services for large-scale LLM workloads on an AI-optimized cloud platform. Define and improve SLIs/SLOs, automate observability with telemetry and performance tuning, and investigate reliability issues using logs and profiling. Partner with AI, platform, and infrastructure teams to optimize training and inference clusters, and help design next-generation distributed systems for AI-first environments.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Design and operate reliable managed AI services for serving and scaling LLM workloads.
  • •Define, measure, and improve SLIs/SLOs to meet performance and reliability targets.
  • •Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters.
  • •Automate observability by building telemetry and performance tuning strategies for latency-sensitive services.
  • •Investigate and resolve reliability issues in distributed AI systems using telemetry, logs, and profiling, and contribute to next-generation distributed systems architecture.

Pay and Benefits

Salary: USD 170,000 - 205,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kEquityPaid ParentalLearning Budget

Key Requirements

  • •Strong software engineering background building production-grade systems beyond scripting or Bash.
  • •Demonstrated experience designing and implementing distributed systems.
  • •SRE mindset: defining/measuring SLIs/SLOs, building monitoring and observability, and driving performance and reliability improvements.
  • •Proficiency in at least one modern programming language: Python, Go, Java, or C++.
  • •Familiarity with Kubernetes or container orchestration platforms.
Experience:Distributed systemsCloud servicesAI infrastructureSRELLM workloads
Skills:CollaborationCommunicationProblem-solvingUrgencyReliability mindset
Tech Stack:PythonGoJavaC++KubernetesContainer orchestrationSLIs/SLOsTelemetryObservabilityLogsProfilingLLMsDistributed systemsTraining clustersInference clusters

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor