Director/Sr. Manager, AI Inference Model Scaling

Cerebras
Sunnyvale
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 12+ yearsEducation: bachelorsSkills: ["Communication","Cross-functional leadership","Organizational planning","Execution","Innovation"]

Lead the Inference Model Scaling organization to enable state-of-the-art foundation models on Cerebras hardware. Define the technical roadmap, organizational strategy, and execution plan for ML compilation, graph optimization, and high-performance kernel development. Hire and grow a globally distributed engineering team, establish engineering standards, and drive design reviews. Partner across compiler, runtime, cloud infrastructure, hardware, product management, and AI research to deliver end-to-end cloud and on-prem inference enablement.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 month ago

Director/Sr. Manager, AI Inference Model Scaling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Lead the Inference Model Scaling organization to enable state-of-the-art foundation models on Cerebras hardware. Define the technical roadmap, organizational strategy, and execution plan for ML compilation, graph optimization, and high-performance kernel development. Hire and grow a globally distributed engineering team, establish engineering standards, and drive design reviews. Partner across compiler, runtime, cloud infrastructure, hardware, product management, and AI research to deliver end-to-end cloud and on-prem inference enablement.
Location: Sunnyvale
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Sr. Manager level

Key Responsibilities

  • •Define the technical roadmap and strategy, including technical direction across teams and engineering leaders.
  • •Lead design reviews and establish engineering standards for inference model scaling.
  • •Hire, mentor, and grow a high-performing engineering team while developing future technical leaders and managers.
  • •Own planning, prioritization, and execution across multiple concurrent initiatives for ML compilation, optimization, and kernel development.
  • •Partner across cloud platform, ML, hardware, product management, and customers to deliver end-to-end service enablement in cloud and on-premise settings.

Key Requirements

  • •BS, MS, or PhD in Computer Science, Computer Engineering or related field.
  • •12+ years building compiler, ML systems, or infrastructure software.
  • •5+ years leading engineering teams.
  • •Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
  • •Experience with Python and C++ and strong understanding of graph compilation and optimization.
Experience:12+ yearsMachine learningCompiler infrastructureDistributed systemsAI infrastructureLLM inference
Education:Bachelor's
Skills:CommunicationCross-functional leadershipOrganizational planningExecutionInnovation
Tech Stack:PythonC++LLVMMLIRXLATVMTorch FXPyTorchJAXTensorFlowONNXWSEWafer-Scale Engine

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn