Senior Site Reliability Engineer

Luma
Redwood City
Workplace: HybridFull timeUSD 168,000 - 252,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Hands-on troubleshooting","Low-level problem solving","Comfort in a fast-paced environment","Less-structured execution","Strong reliability mindset"]

Own production GPU clusters that support training and inference across AWS and OCI, ensuring high availability and performance. Triage and resolve complex GPU, networking (InfiniBand/RDMA), and kernel-level failures, sometimes debugging directly with NVIDIA. Deeply tune Linux/OS performance, build automation in Python/Go/Bash, and participate in re-architecture for next-level scale while strengthening infrastructure security practices for SOC 2 and ISO.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Luma
Luma
1 week ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own production GPU clusters that support training and inference across AWS and OCI, ensuring high availability and performance. Triage and resolve complex GPU, networking (InfiniBand/RDMA), and kernel-level failures, sometimes debugging directly with NVIDIA. Deeply tune Linux/OS performance, build automation in Python/Go/Bash, and participate in re-architecture for next-level scale while strengthening infrastructure security practices for SOC 2 and ISO.
Location: Redwood City
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
  • •Participate in critical re-architecture sessions to redesign systems for higher efficiency and scale.
  • •Tune Linux performance deeply at the OS and kernel level.
  • •Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
  • •Serve as final escalation for the hardest GPU, networking, and system failures, working with vendors like NVIDIA.

Pay and Benefits

Salary: USD 168,000 - 252,000 annually
Equity and Bonus:Equity

Key Requirements

  • •5+ years as an SRE, production, or infrastructure engineer in fast-paced, large-scale environments.
  • •Deep, hands-on Linux expertise, including containerized systems and low-level performance debugging.
  • •Working experience with Terraform, Airflow, and Ray.
  • •Strong experience with AWS or OCI.
  • •Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
Experience:AI/MLInfrastructureSRECloud computing
Skills:Hands-on troubleshootingLow-level problem solvingComfort in a fast-paced environmentLess-structured executionStrong reliability mindset
Certifications:SOC 2 Type 1SOC 2 Type 2ISO
Tech Stack:LinuxPythonGoBashTerraformAirflowRayAWSOCIInfiniBandRDMARoCENVIDIAAMDDCGMROCmKubernetesSOC 2ISOSOC 2 Type 1

Company Brief

Luma
Develops AI-powered tools for capturing, editing, and rendering high-quality 3D scenes from photos and videos, enabling creators to generate photorealistic 3D assets and spatial experiences.
Industry: AR/VR & Spatial Computing
Website