Staff Network Production Engineer, Operations

Crusoe
San Francisco
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Experience: 8+ yearsEducation: bachelorsSkills: ["Communication","Leadership","Problem-solving","Mentoring"]

Seasoned Staff Network Operations Engineer responsible for production reliability across Crusoe's global edge, backbone, data center, and GPU cluster networks. Leads incident response, drives root cause analysis, enhances observability, builds automation in Python, and mentors peers to uphold high-availability of AI infrastructure at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
3 months ago

Staff Network Production Engineer, Operations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Seasoned Staff Network Operations Engineer responsible for production reliability across Crusoe's global edge, backbone, data center, and GPU cluster networks. Leads incident response, drives root cause analysis, enhances observability, builds automation in Python, and mentors peers to uphold high-availability of AI infrastructure at scale.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Production Reliability: Own uptime across Crusoe's global edge, backbone, data center, and GPU cluster network, directly supporting AI workloads at scale.
  • •Incident Response: Lead and contribute to end-to-end response for high-severity network events, including mitigation, stakeholder communication, and postmortem documentation.
  • •Root Cause Analysis: Drive RCAs for production incidents, identify systemic issues, and author remediation plans tracked through to closure.
  • •Observability Improvements: Contribute to and improve Crusoe's network monitoring stack using streaming telemetry, SNMP, NetFlow, and tools such as Kentik, Grafana, Prometheus, and ThousandEyes.
  • •Operational Automation: Write Python-based tooling to reduce toil, automate common remediation workflows, and accelerate mean time to resolution.

Pay and Benefits

Perks:Health InsuranceDentalVision401kEquityRsusParental LeaveLife InsuranceDisabilityTravel Allowance

Key Requirements

  • •8+ years of production network engineering experience with a focus on operations, incident response, and reliability in large-scale or internet-scale environments.
  • •Hands-on experience with observability and monitoring tools including streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes.
  • •Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including familiarity with PFC, ECN, and DCQCN tuning.
  • •Expert hands-on knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments.
  • •Proficiency with Arista (EOS) and Juniper (Junos) platforms in leaf-spine CLOS architectures across multi-vendor environments.
Experience:8+ yearsCloud networkingAI infrastructureData centerGPU clusters
Education:Bachelor's
Skills:CommunicationLeadershipProblem-solvingMentoring
Tech Stack:PythonGrafanaPrometheusThousandEyesSNMPNetFlowSFlowKentikArista EOSJuniper JunoBGPEVPN-VXLANIS-ISOSPFMPLSQoSTCP/IPRDMARoCEPFC

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor