Site Reliability Engineer

X AI
Tennessee, Memphis, Mississippi
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Incident leadership","Calm technical leadership","Problem-solving","Data-driven approach","Cross-functional collaboration"]

Design and operate campus-level reliability systems by defining what is monitored, trusted, and alerted. Lead SEV-class incidents with technical command and NOC coordination, run blameless postmortems, and drive corrective actions to closure. Own monitoring signal quality and alert hygiene, set error budgets and availability objectives, and partner on cross-discipline projects spanning compute, network, storage, power, and cooling. Build playbooks, run game days, and support 24/7 operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
2 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Design and operate campus-level reliability systems by defining what is monitored, trusted, and alerted. Lead SEV-class incidents with technical command and NOC coordination, run blameless postmortems, and drive corrective actions to closure. Own monitoring signal quality and alert hygiene, set error budgets and availability objectives, and partner on cross-discipline projects spanning compute, network, storage, power, and cooling. Build playbooks, run game days, and support 24/7 operations.
Location: Tennessee, Memphis, Mississippi
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own monitoring architecture and signal quality, including what to alert on, suppress, and trust; use NOC noise feedback to improve alerting.
  • •Provide SEV command support with technical incident leadership, bridge coordination with the NOC, and maintain timeline/severity hygiene.
  • •Run blameless postmortems and drive corrective actions to closure (not filed).
  • •Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • •Build and maintain playbooks, run game days, keep dependency maps current, and define error budgets and availability objectives; participate in on-call rotations.

Key Requirements

  • •Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • •5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
  • •Proven large-scale incident command experience with calm technical leadership.
  • •Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • •Experience writing and operating playbooks/runbooks with a 24/7 operations or NOC partner.
Experience:5+ yearsHigh-performance computingData centerProduction operations
Education:Bachelor's in Systems Engineering, Computer Science, Electrical Engineering, or related field
Skills:Incident leadershipCalm technical leadershipProblem-solvingData-driven approachCross-functional collaboration
Languages:English
Tech Stack:PythonBashCC++JavaGoRust

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website