Senior Site Reliability Engineer I

Braze
Toronto
Workplace: OnsiteFull timeCAD 153,815 - 277,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Documentation","Collaboration","Systems thinking"]

Own the ingress and API-ingestion infrastructure that keeps Braze’s distributed services highly available and scalable. Lead NGINX and Kubernetes ingress design, scaling routines using RED metrics and HPA, and capacity planning for enterprise-grade SLAs. Partner with product engineering teams on resilient architectures, manage SLIs/SLOs and error budgets, and deliver on-call incident response with runbooks and blameless RCAs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Braze
Braze
2 days ago

Senior Site Reliability Engineer I

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Own the ingress and API-ingestion infrastructure that keeps Braze’s distributed services highly available and scalable. Lead NGINX and Kubernetes ingress design, scaling routines using RED metrics and HPA, and capacity planning for enterprise-grade SLAs. Partner with product engineering teams on resilient architectures, manage SLIs/SLOs and error budgets, and deliver on-call incident response with runbooks and blameless RCAs.
Location: Toronto
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead NGINX and Kubernetes ingress infrastructure, configuring, tuning, and operating high-performance ingress controller layers for API traffic.
  • •Own and expand automated scaling routines using RED metrics, Horizontal Pod Autoscalers (HPA), and customized scaling policies.
  • •Partner with product engineering teams to translate feature requirements into resilient, highly available, and scalable technology stacks.
  • •Establish and manage SLIs, SLOs, and error budgets for API services to balance deployment speed with stability.
  • •Participate in PagerDuty on-call rotation, maintain runbooks, and lead blameless root-cause analysis for incidents.

Pay and Benefits

Salary: CAD 153,815 - 277,000 annually
Perks:RetirementEquityPaid LeaveHealth InsuranceDentalVisionLife InsuranceDisability InsuranceParental LeaveLearning Budget

Key Requirements

  • •5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
  • •Deep hands-on expertise configuring, troubleshooting, and operating high-performance NGINX proxying, routing, and ingress controllers under heavy traffic.
  • •In-depth, hands-on proficiency with Kubernetes administration, cluster networking, container orchestration, cluster scheduling, and deployments.
  • •Excellent Linux/Unix fundamentals including disk I/O, memory allocation, TCP/IP networking, and process management.
  • •Strong programming/scripting skills (Ruby and/or Go preferred) and experience with infrastructure as code (e.g., Terraform, Ansible, Chef).
Skills:DocumentationCollaborationSystems thinking
Languages:English
Tech Stack:NGINXKubernetesRubyGoPythonJavaLinuxTerraformAnsibleChefPagerDutyHPARED metricsPrometheusGrafanaDatadogRedisKafkaPostgresMongoDB

Company Brief

Braze
Provides a customer engagement platform that helps brands create personalized messaging and lifecycle campaigns across mobile, web, email, and other channels to drive retention, engagement, and revenue.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 100M to 250M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
WebsiteLinkedIn