Senior Site Reliability Engineer
Redwood City
Workplace: HybridFull timeUSD 168,000 - 252,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Hands-on troubleshooting","Low-level problem solving","Comfort in a fast-paced environment","Less-structured execution","Strong reliability mindset"]Own production GPU clusters that support training and inference across AWS and OCI, ensuring high availability and performance. Triage and resolve complex GPU, networking (InfiniBand/RDMA), and kernel-level failures, sometimes debugging directly with NVIDIA. Deeply tune Linux/OS performance, build automation in Python/Go/Bash, and participate in re-architecture for next-level scale while strengthening infrastructure security practices for SOC 2 and ISO.
Loading
Loading job details...
Preparing the role view and application actions.

