Senior Site Reliability Engineer I

Braze
São Paulo
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Collaboration","Documentation","Fast delivery","Problem-solving","Bias toward action"]

Own site reliability for internal-facing services by partnering with engineering teams to build scalable, reliable infrastructure and improve monitoring, alerting, and incident response. Develop infrastructure as code and deployment pipelines using Chef, Terraform, Kubernetes, and Docker. Debug reliability issues across stack layers, ensure enterprise-grade SLAs, and use on-call rotations plus continuous retrospectives to drive automation and operational improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Braze
Braze
1 day ago

Senior Site Reliability Engineer I

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Own site reliability for internal-facing services by partnering with engineering teams to build scalable, reliable infrastructure and improve monitoring, alerting, and incident response. Develop infrastructure as code and deployment pipelines using Chef, Terraform, Kubernetes, and Docker. Debug reliability issues across stack layers, ensure enterprise-grade SLAs, and use on-call rotations plus continuous retrospectives to drive automation and operational improvements.
Location: São Paulo
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Partner with engineering teams to architect scalable, reliable products using infrastructure platforms.
  • •Debug reliability and scalability issues across the full stack, including products built on infrastructure platforms.
  • •Design monitoring and alerting based on symptoms rather than outages to prevent and reduce incidents.
  • •Develop internal platform infrastructure using Infrastructure as Code (Chef, Terraform) and Kubernetes, plus deployment pipelines with Docker/Kubernetes.
  • •Manage incidents through PagerDuty rotation: respond to availability incidents, prevent issues using on-call, and run retrospectives to turn lessons into system improvements.

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsurancePaid LeaveRetirementEquityLearning Budget

Key Requirements

  • •3+ years as a Software, DevOps, or Site Reliability Engineer.
  • •Strong systems thinking across interfaces, boundaries, edge cases, and failure modes.
  • •Experience with Linux and Unix Shell, plus strong programming skills in Ruby and/or Go.
  • •Experience with Docker, Kubernetes, and Terraform (or similar IaC technologies).
  • •Experience with MongoDB, Redis, Kafka, Postgres (or similar data technologies).
Experience:3+ yearsInfrastructureDistributed systemsDevOps
Skills:CollaborationDocumentationFast deliveryProblem-solvingBias toward action
Tech Stack:Ruby on RailsRubyGoMongoDBRedisKafkaPostgresChefTerraformKubernetesDockerPagerDutyLinuxUnix ShellInfrastructure as code

Company Brief

Braze
Provides a customer engagement platform that helps brands create personalized messaging and lifecycle campaigns across mobile, web, email, and other channels to drive retention, engagement, and revenue.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 100M to 250M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
WebsiteLinkedIn