Senior Site Reliability Engineer I

Braze
Vancouver
Full timeCAD 153,815 - 277,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Documentation","Collaboration","Systems thinking","Proactive incident management","Root-cause analysis"]

Own and scale the ingress infrastructure that powers how the internet communicates with Braze. In this embedded SRE role, lead NGINX and Kubernetes ingress operations, architect resilient ingress fleets, and tune automated scaling using RED metrics and HPA. Partner with product engineering teams to define SLIs/SLOs, drive systems design and capacity planning, and deliver incident response excellence through on-call, runbook improvements, and blameless RCA.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Braze
Braze
2 days ago

Senior Site Reliability Engineer I

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Own and scale the ingress infrastructure that powers how the internet communicates with Braze. In this embedded SRE role, lead NGINX and Kubernetes ingress operations, architect resilient ingress fleets, and tune automated scaling using RED metrics and HPA. Partner with product engineering teams to define SLIs/SLOs, drive systems design and capacity planning, and deliver incident response excellence through on-call, runbook improvements, and blameless RCA.
Location: Vancouver
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead NGINX and Kubernetes ingress infrastructure for high-throughput, real-time API ingestion.
  • •Architect and operate ingress fleets, including configuring, tuning, and operating ingress controller layers.
  • •Own and expand automated scaling routines using RED metrics, HPA, and customized scaling policies.
  • •Partner with product engineering teams to translate requirements into resilient, highly available technology stacks and define SLIs/SLOs and error budgets.
  • •Participate in PagerDuty on-call, maintain runbooks, and lead blameless RCA and retrospectives to drive permanent improvements.

Pay and Benefits

Salary: CAD 153,815 - 277,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeaveLearning BudgetLife InsuranceDisabilityRetirementEquityRsusParental Leave

Key Requirements

  • •5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
  • •Deep hands-on experience configuring, troubleshooting, and operating high-performance NGINX ingress/proxy layers under heavy traffic.
  • •In-depth Kubernetes administration skills, including cluster networking, orchestration, scheduling, and container deployment.
  • •Strong Linux/Unix fundamentals (disk I/O, memory, TCP/IP networking, and process management).
  • •Strong programming/scripting ability (Ruby and/or Go preferred; Python/Java acceptable) and experience with Infrastructure as Code (e.g., Terraform, Ansible, Chef).
Experience:5+ yearsSREDevOpsDistributed systems
Skills:DocumentationCollaborationSystems thinkingProactive incident managementRoot-cause analysis
Tech Stack:NGINXKubernetesLinuxRubyGoPythonJavaTerraformAnsibleChefPagerDutyRED metricsHorizontal Pod Autoscalers (HPA)PrometheusGrafanaDatadogRedisKafkaPostgresMongoDB

Company Brief

Braze
Provides a customer engagement platform that helps brands create personalized messaging and lifecycle campaigns across mobile, web, email, and other channels to drive retention, engagement, and revenue.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 100M to 250M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
WebsiteLinkedIn