Staff Engineer, Datacenter Server Lifecycle

Anthropic
Sydney
Workplace: HybridFull timeFunction: Hospitality & Food ServiceExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Stakeholder management","Consensus building","Navigating ambiguity","Cross-functional collaboration"]

Own the end-to-end lifecycle of datacenter servers across global facilities, from provisioning and deployment through steady-state operation, maintenance, refresh, and decommissioning. Build automation and operational procedures for fleet-wide lifecycle events at tens-of-thousands scale, while partnering with Infrastructure Security to enforce trusted compute standards. Collaborate with Networking to ensure connectivity across sites and develop tooling to track machine health, configuration, and operational status.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
1 month ago

Staff Engineer, Datacenter Server Lifecycle

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Own the end-to-end lifecycle of datacenter servers across global facilities, from provisioning and deployment through steady-state operation, maintenance, refresh, and decommissioning. Build automation and operational procedures for fleet-wide lifecycle events at tens-of-thousands scale, while partnering with Infrastructure Security to enforce trusted compute standards. Collaborate with Networking to ensure connectivity across sites and develop tooling to track machine health, configuration, and operational status.
Location: Sydney
Workplace: Hybrid
Employment Type: Full time
Job Function: Hospitality & Food Service
Seniority: Sr. Manager level

Key Responsibilities

  • •Build automation to support datacenters containing tens of thousands of servers.
  • •Define and own end-to-end server lifecycle strategy, including provisioning/deployment through operation, maintenance, refresh, and decommissioning.
  • •Partner with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle.
  • •Work with Networking to ensure end-to-end connectivity across all sites.
  • •Build and maintain tooling to track machine health, configuration, and operational status across the datacenter fleet.
Travel: Low travel

Pay and Benefits

Perks:Paid LeaveParental LeaveEquity

Key Requirements

  • •Hands-on experience with server hardware, including rack deployment, cabling, troubleshooting, and understanding failure modes at scale.
  • •End-to-end understanding of hardware lifecycle management, including asset tracking, provisioning workflows, maintenance scheduling, and decommissioning practices.
  • •Proficiency in at least one programming language such as Python, Rust, Go, or Java.
  • •Working knowledge of modern cloud infrastructure, including Kubernetes and large-scale cloud providers like AWS, Azure, and GCP.
  • •Ability to communicate clearly and build consensus with a wide range of stakeholders, while navigating ambiguity on cross-functional problems.
Experience:8+ yearsDatacenter infrastructureCloud infrastructureFleet managementAI infrastructure
Education:Bachelor's
Skills:CommunicationStakeholder managementConsensus buildingNavigating ambiguityCross-functional collaboration
Languages:English
Tech Stack:PythonRustGoJavaKubernetesAWSAzureGCPLinuxBootNixOSLinuxTPMSecure bootHardware attestationFirmware verificationNVIDIA A100NVIDIA H100Google TPUsAWS Trainium

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn