Staff Machine Learning Engineer, ML Platform

Braze
San Francisco, New York
Workplace: HybridFull timeUSD 184,000 - 314,000 annuallyFunction: Data Science & Machine LearningExperience: 8+ yearsSkills: ["Autonomy","Accountability","Communication","Technical leadership","Mentorship"]

Own Braze’s ML platform powering large-scale training and real-time, multi-region model serving for personalized customer engagement. Lead complex infrastructure initiatives end-to-end—replatforming orchestration, deployment and cloud identity, and improving reliability and cost. Set the technical vision and quality bar for training, deployment, serving, and observability, drive incident response, mentor senior engineers and data scientists, and coordinate across teams to deliver production impact.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Braze
Braze
4 days ago

Staff Machine Learning Engineer, ML Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own Braze’s ML platform powering large-scale training and real-time, multi-region model serving for personalized customer engagement. Lead complex infrastructure initiatives end-to-end—replatforming orchestration, deployment and cloud identity, and improving reliability and cost. Set the technical vision and quality bar for training, deployment, serving, and observability, drive incident response, mentor senior engineers and data scientists, and coordinate across teams to deliver production impact.
Location: San Francisco, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Identify and drive production ML platform initiatives, including replatforming queueing/orchestration, deployment and cloud identity, and retiring infrastructure generations.
  • •Design, build, and ship ML platform infrastructure at high velocity, including multi-region model serving fleets and CI/deployment tooling.
  • •Own the technical vision and production quality bar for training, deployment, serving, and observability, including leading incident response for ML systems.
  • •Drive reliability and cost work at scale and ensure operational efficiency across global production workloads.
  • •Coordinate cross-team initiatives, raise engineering quality through design/code reviews and production readiness, and mentor senior engineers and data scientists.

Pay and Benefits

Salary: USD 184,000 - 314,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRsusRetirementPaid LeaveParental LeaveLearning Budget

Key Requirements

  • •8+ years building and operating distributed systems in production, with depth in deployment and operations.
  • •Hands-on experience with ML workloads in production, including training pipelines, model serving, or ML platform tooling.
  • •Owned direction for a team and led multi-quarter initiatives across team boundaries, with high personal output.
  • •Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management and networking.
  • •Owned CI/CD and infrastructure as code, and run systems under production load with scale and reliability.
Experience:8+ years
Skills:AutonomyAccountabilityCommunicationTechnical leadershipMentorship
Languages:English
Tech Stack:KubernetesCI/CDInfrastructure as codePythonRuby on RailsMongoDBRedisCeleryRabbitMQKafkaRayMLflow

Company Brief

Braze
Provides a customer engagement platform that helps brands create personalized messaging and lifecycle campaigns across mobile, web, email, and other channels to drive retention, engagement, and revenue.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 100M to 250M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
WebsiteLinkedIn