Senior Machine Learning Systems Engineer

Reddit
United States
Workplace: RemoteFull timeFunction: IT Operations (Systems/Network Admin)Experience: 5+ yearsSkills: ["Communication","Collaboration","Problem-solving"]

Lead development of a platform for large-scale ML models, shaping end-to-end ML lifecycle patterns (MLOps) and a graph ML codebase. Focus on performance, scalability, and cost-efficiency in a distributed training environment, collaborating with ML engineers to optimize training, data processing, and model deployment within Reddit’s ML Infrastructure domain.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reddit
Reddit
4 months ago

Senior Machine Learning Systems Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Lead development of a platform for large-scale ML models, shaping end-to-end ML lifecycle patterns (MLOps) and a graph ML codebase. Focus on performance, scalability, and cost-efficiency in a distributed training environment, collaborating with ML engineers to optimize training, data processing, and model deployment within Reddit’s ML Infrastructure domain.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Design end-to-end ML lifecycle patterns (MLOps) to boost velocity of development for ML engineers, including data preparation, model management, experiment tracking, and more
  • •Zero-to-one development and support of a graph ML codebase and platform that abstracts away common patterns and enables greater model scalability and iteration
  • •Collaborate with ML engineers on performance tuning, including improving model training time, efficiency, and GPU training costs in a large, distributed ML training environment
  • •Optimize batch data processing within a data warehouse and with tools such as Apache Beam, Apache Spark, Ray Data, and more
  • •Architect pipelines to build and maintain massive graph data structures on the order of billions of nodes and tens of billions of edges

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceVision401kParonal LeavePaid Leave

Key Requirements

  • •5+ years of experience in ML infrastructure, including model training and model deployments
  • •Hands-on experience with ML optimization, including memory and GPU profiling
  • •Deep experience with cloud-based technologies for supporting an ML platform, including tools like GCP BigQuery, Google Cloud Storage, infrastructure-as-code (Terraform)
  • •Hands-on experience administering and integrating MLOps tools for experiment tracking, model serving, and model registries (e.g. MLflow or Wandb)
  • •Proficiency with programming languages and frameworks of ML (Python, PyTorch, TensorFlow)
Experience:5+ yearsMachine learningMl platform
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonPyTorchTensorFlowGCP BigQueryGoogle Cloud StorageTerraformMLflowWandBRayKubernetesApache BeamApache SparkNeo4jJanusGraphTigerGraphPyTorch GeometricDeep Graph LibraryGPUDistributed training

Company Brief

Reddit
Operates Reddit, a large online community and discussion platform where users submit content, comment, and vote across topic-based communities (subreddits); monetizes via advertising, premium subscriptions, and awards.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2005
WebsiteLinkedIn