Member of Technical Staff - Reliability Engineering

Fireworks AI
San Mateo, New York
Workplace: HybridFull timeUSD 240,000 - 290,000 annuallyFunction: Data Analytics & Business IntelligenceExperience: 5+ yearsEducation: mastersSkills: ["Ownership","Troubleshooting","Influence without authority","Cross-boundary collaboration","Attention to reliability details"]

Ensure Fireworks’ AI infrastructure runs dependably as it scales by defining reliability standards (SLOs, error budgets, production readiness) and owning the reliability toolchain. Build and improve observability, alerting, failure testing, incident management, and self-healing automation. Partner across cloud, inference/training, and product teams to fix failure modes at system seams, reduce toil, and deliver reliable customer experiences under load.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fireworks AI
Fireworks AI
1 day ago

Member of Technical Staff - Reliability Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Ensure Fireworks’ AI infrastructure runs dependably as it scales by defining reliability standards (SLOs, error budgets, production readiness) and owning the reliability toolchain. Build and improve observability, alerting, failure testing, incident management, and self-healing automation. Partner across cloud, inference/training, and product teams to fix failure modes at system seams, reduce toil, and deliver reliable customer experiences under load.
Location: San Mateo, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Mid level

Key Responsibilities

  • •Define reliability standards including SLOs, error budgets, and production readiness criteria using real telemetry.
  • •Own the reliability toolchain: logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and investigation tooling.
  • •Ensure customer experience reliability by assigning ownership and fixes for per-service failures that may not breach system SLOs.
  • •Identify and drive improvements for failure modes at system seams (e.g., retries, timeouts, unmapped dependencies).
  • •Run incident management by coordinating production issues, leading blameless postmortems, and tracking follow-ups to completion.

Pay and Benefits

Salary: USD 240,000 - 290,000 annually
Equity and Bonus:Equity

Key Requirements

  • •5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC).
  • •5+ years in Python, Go, C++, or Rust, building production-grade systems and tools.
  • •Operate and debug Kubernetes, Terraform, and Docker in high-throughput production environments.
  • •Demonstrated distributed systems experience (high-throughput control planes, microservices, or multi-region setups).
  • •Reliability fundamentals including fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture; bachelor’s or master’s in CS/CE (or equivalent).
Experience:5+ yearsAI infrastructureCloud-nativeDistributed systemsOpen modelsInference servingTraining infrastructure
Education:Master's in Computer Science, Computer Engineering
Skills:OwnershipTroubleshootingInfluence without authorityCross-boundary collaborationAttention to reliability details
Tech Stack:LinuxPythonGoC++RustTCP/IPHTTPGRPCKubernetesTerraformDockerPrometheusGrafanaOpenTelemetryGPUMicroservicesMulti-regionObservabilityStructured loggingMetrics

Company Brief

Fireworks AI
Develops AI-driven tools to generate and optimize visual marketing content for brands and creators, automating production of short-form videos and multimedia assets for social platforms to improve engagement and scale creative workflows.
Industry: SaaS
Website