Senior Machine Learning Systems Engineer, Ads ML Experience Platform

Reddit
United States
Workplace: RemoteFull timeUSD 216,700 - 303,400 annuallyFunction: IT Operations (Systems/Network Admin)Experience: 5+ yearsSkills: ["Collaboration"]

Design and build large-scale offline ML experimentation platforms and production training orchestration frameworks that accelerate the Ads ML lifecycle from experimentation through evaluation and autonomous deployment. Create infrastructure for experiment tracking, metadata management, lineage, artifact versioning, and model registries to ensure reproducibility. Partner with ML engineers and researchers to improve iteration speed, and build an agentic AI execution platform with multi-agent orchestration and scalable workflow infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reddit
Reddit
1 day ago

Senior Machine Learning Systems Engineer, Ads ML Experience Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Design and build large-scale offline ML experimentation platforms and production training orchestration frameworks that accelerate the Ads ML lifecycle from experimentation through evaluation and autonomous deployment. Create infrastructure for experiment tracking, metadata management, lineage, artifact versioning, and model registries to ensure reproducibility. Partner with ML engineers and researchers to improve iteration speed, and build an agentic AI execution platform with multi-agent orchestration and scalable workflow infrastructure.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Design and build large-scale offline ML experimentation platforms for reproducible research, evaluation, and promotion workflows.
  • •Develop production-grade training orchestration frameworks for distributed training, hyperparameter optimization, evaluation, and automated retraining.
  • •Build infrastructure for experiment tracking, metadata management, lineage, artifact versioning, model registries, and reproducibility.
  • •Partner with ML engineers and researchers to improve experimentation velocity and operational efficiency.
  • •Design and build an agentic AI execution platform for autonomous and human-in-the-loop workflows, including multi-agent orchestration and scalable workflow infrastructure.

Pay and Benefits

Salary: USD 216,700 - 303,400 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid LeaveParental Leave

Key Requirements

  • •5+ years in infrastructure/platform engineering or large-scale distributed systems.
  • •2+ years building and operating production ML infrastructure, developer SDKs, platform APIs, or self-service AI tooling.
  • •Experience building workflow orchestration systems, developer platforms, or large-scale automation frameworks.
  • •Experience with distributed data processing systems such as Spark, Flink, Ray, or equivalents.
  • •Experience with orchestration/workflow technologies such as Kubeflow, Argo, Airflow, or similar frameworks.
Experience:5+ yearsMachine learningDistributed systemsAgentic AI
Skills:Collaboration
Languages:English
Tech Stack:SparkFlinkRayKubeflowArgoAirflowMCPA2A

Company Brief

Reddit
Operates Reddit, a large online community and discussion platform where users submit content, comment, and vote across topic-based communities (subreddits); monetizes via advertising, premium subscriptions, and awards.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2005
WebsiteLinkedIn