Sr./Staff TPM - Inference Capacity

Cerebras
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: Program & Project Management (PMO)Experience: 5+ yearsSkills: ["Cross-functional collaboration","Capacity planning","Risk mitigation","Data fluency","Process improvement"]

Lead capacity planning and fleet strategy for the Inference Service organization, translating customer contracts and model launches into a rolling capacity plan across clusters. Own forecasting and utilization reporting, support new datacenter bring-up for production readiness, and manage allocation and cluster placement with SRE and product teams. Drive adoption of internal capacity management tools, proactively track capacity risks and bottlenecks, and maintain Jira EPICs and Confluence pages for cross-team execution transparency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 months ago

Sr./Staff TPM - Inference Capacity

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead capacity planning and fleet strategy for the Inference Service organization, translating customer contracts and model launches into a rolling capacity plan across clusters. Own forecasting and utilization reporting, support new datacenter bring-up for production readiness, and manage allocation and cluster placement with SRE and product teams. Drive adoption of internal capacity management tools, proactively track capacity risks and bottlenecks, and maintain Jira EPICs and Confluence pages for cross-team execution transparency.
Location: Sunnyvale, Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Program & Project Management (PMO)
Seniority: Mid level

Key Responsibilities

  • •Build and maintain the 6/12/26-week rolling capacity model across clusters and reconcile forecasts against weekly actuals.
  • •Collaborate on new datacenter capacity bring-up to ensure production readiness, on-time delivery, and quality.
  • •Run weekly capacity reviews and decide model placement and re-balancing for tenants, clusters, and freezes; report utilization and forecast.
  • •Drive stakeholder adoption of the in-house capacity planning and allocation tool, including UAT, issue resolution, pilot testing, and deployment.
  • •Proactively identify and mitigate capacity bottlenecks and SLA-drop risks, and lead incident tracking and postmortems; maintain Jira EPICs and Confluence pages.

Key Requirements

  • •5+ years of technical program management (TPM), technical program management, or product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning.
  • •Experience leading large cross-functional programs across Engineering, Product, and Operations.
  • •Comfort with the inference serving stack including model replicas, batching, prefill/decode, KV cache, and accelerator scheduling.
  • •Strong data fluency with SQL and Grafana, plus basic Python or Flux to pull numbers independently.
  • •Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, and Trainium.
Experience:5+ yearsCloud infrastructureML servingHyperscaler capacity planningAI accelerator fleets
Skills:Cross-functional collaborationCapacity planningRisk mitigationData fluencyProcess improvement
Tech Stack:SQLGrafanaPythonFluxJiraConfluence

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn