Staff Machine Learning Systems Engineer

Reddit
United States
Workplace: RemoteFull timeUSD 230,000 - 322,000 annuallyFunction: IT Operations (Systems/Network Admin)Experience: 8+ yearsSkills: ["Communication","Collaboration"]

Lead development of a platform for large scale ML models, focusing on end-to-end ML lifecycle patterns, graph ML infrastructure, and performance optimization in a distributed environment. You will collaborate with ML engineers to improve training efficiency, manage data pipelines, and scale graph data structures, enabling faster model iteration and deployment across Reddit’s ML infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reddit
Reddit
4 months ago

Staff Machine Learning Systems Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Lead development of a platform for large scale ML models, focusing on end-to-end ML lifecycle patterns, graph ML infrastructure, and performance optimization in a distributed environment. You will collaborate with ML engineers to improve training efficiency, manage data pipelines, and scale graph data structures, enabling faster model iteration and deployment across Reddit’s ML infrastructure.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Design end-to-end model lifecycle patterns (MLOps) to boost velocity of development for ML engineers, including data preparation, model management, experiment tracking, and more.
  • •Zero-to-one development and support of a graph ML codebase and platform that abstracts away common patterns and enables greater model scalability and iteration.
  • •Collaborate with ML engineers on performance tuning, including improving model training time, efficiency, and GPU training costs in a large, distributed ML training environment.
  • •Optimize batch data processing within a data warehouse and with tools such as Apache Beam, Apache Spark, Ray Data, and more.
  • •Architect pipelines to build and maintain massive graph data structures on the order of billions of nodes and tens of billions of edges.

Pay and Benefits

Salary: USD 230,000 - 322,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kParental LeavePaid Leave

Key Requirements

  • •8+ years of experience in ML infrastructure, including model training and model deployments.
  • •Memory and GPU profiling experience for ML optimization.
  • •Cloud-based ML platform experience, including tools like GCP BigQuery and Google Cloud Storage, and infrastructure-as-code (Terraform).
  • •Experience administering and integrating MLOps tools for experiment tracking, model serving, and model registries (e.g., MLflow or Wandb).
  • •Proficiency with Python, PyTorch, TensorFlow, and distributed training frameworks (Ray, Kubernetes).
Experience:8+ yearsMachine learning infrastructureGraph MLDistributed training
Skills:CommunicationCollaboration
Languages:English
Tech Stack:PythonPyTorchTensorFlowApache BeamApache SparkRayKubernetesGCP BigQueryGoogle Cloud StorageTerraformMLflowWandbNeo4jJanusGraphTigerGraph

Company Brief

Reddit
Operates Reddit, a large online community and discussion platform where users submit content, comment, and vote across topic-based communities (subreddits); monetizes via advertising, premium subscriptions, and awards.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2005
WebsiteLinkedIn