Engineering Manager, AI Platform - Managed AI

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull timeUSD 215,000 - 260,000 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["People leadership","Mentorship","Problem-solving","Communication","Collaboration"]

Lead and scale a team of engineers building a next-generation platform for the full lifecycle of large language models (LLMs). Own the architecture and delivery of scalable, fault-tolerant AI services, including task queues, model management systems, and cost-aware scheduling. Partner with Product, Infrastructure, and GTM stakeholders to shape the engineering roadmap, influence strategic discussions, and ensure the platform can support massive API request throughput.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
2 days ago

Engineering Manager, AI Platform - Managed AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Lead and scale a team of engineers building a next-generation platform for the full lifecycle of large language models (LLMs). Own the architecture and delivery of scalable, fault-tolerant AI services, including task queues, model management systems, and cost-aware scheduling. Partner with Product, Infrastructure, and GTM stakeholders to shape the engineering roadmap, influence strategic discussions, and ensure the platform can support massive API request throughput.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Manager level

Key Responsibilities

  • •Lead, mentor, and grow a team of engineers on the Managed AI platform.
  • •Define goals and drive accountability for the AI roadmap in partnership with leadership.
  • •Oversee architecture and development of core AI services (fault-tolerant task queues, model management, cost-aware scheduling).
  • •Ensure scalable, fault-tolerant systems that can handle very high API request throughput.
  • •Collaborate cross-functionally with Product, Infrastructure, and GTM stakeholders to drive delivery and adoption.

Pay and Benefits

Salary: USD 215,000 - 260,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRsus401kPaid LeaveParental LeaveHsaLife Insurance

Key Requirements

  • •5+ years managing/leading high-performing engineering teams.
  • •Hands-on experience with distributed and concurrent systems or AI infrastructure.
  • •Deep knowledge of cloud-native environments, container orchestration, and SOAs.
  • •Proficiency in Python/GoLang/Rust.
  • •Experience with Kubernetes, gRPC, and observability stacks.
Experience:5+ yearsAI infrastructureLLM platformsCloud-nativeDistributed systemsStartup environment
Skills:People leadershipMentorshipProblem-solvingCommunicationCollaboration
Tech Stack:PythonGoLangRustKubernetesGRPCObservabilityVLLMHugging FaceTritonCPUGPULLMTask queuesModel management systemsCost-aware schedulingContainersSOA

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor