Senior Software Engineer, Machine Learning Infrastructure

Match Group
United States
Workplace: HybridFull timeUSD 190,000 - 246,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Technical guidance","Mentorship","Cross-functional collaboration","Technical documentation","Problem-solving"]

Design and operate scalable machine learning infrastructure that powers experimentation, training, deployment, and monitoring for large-scale datasets. Build robust data processing and moderation pipelines integrated with trust and safety workflows, and develop ML platform systems with distributed data technologies. Implement APIs, observability, model evaluation/A-B testing frameworks, and CI/CD/GitOps automation to ensure reliable, cost-efficient ML services. Guide engineers and help evaluate candidates while shaping ML lifecycle infrastructure, including LLM workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Match Group
Match Group
2 days ago

Senior Software Engineer, Machine Learning Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Design and operate scalable machine learning infrastructure that powers experimentation, training, deployment, and monitoring for large-scale datasets. Build robust data processing and moderation pipelines integrated with trust and safety workflows, and develop ML platform systems with distributed data technologies. Implement APIs, observability, model evaluation/A-B testing frameworks, and CI/CD/GitOps automation to ensure reliable, cost-efficient ML services. Guide engineers and help evaluate candidates while shaping ML lifecycle infrastructure, including LLM workloads.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain scalable ML infrastructure for experimentation, training, deployment, and monitoring at large scale.
  • •Develop robust data processing and moderation pipelines integrated with trust and safety workflows, and optimize compute/storage for reliability, scalability, and cost efficiency.
  • •Create and maintain internal ML platform services and integrations via APIs (REST, gRPC, GraphQL).
  • •Implement deployment monitoring, performance observability, and model evaluation/validation/quality assurance including A/B testing and automated evaluation systems.
  • •Lead platform engineering efforts end-to-end, including model training/serving/feature stores/evaluation systems, CI/CD and GitOps automation, mentoring engineers, and supporting hiring through technical interviews.
Travel: Low travel

Pay and Benefits

Salary: USD 190,000 - 246,000 annually

Key Requirements

  • •Bachelor’s degree (or U.S. equivalent) in Computer Science, Computer Engineering, or a related field, plus 5 years of professional experience in Machine Learning Engineering, Site Reliability Engineering, or ML infrastructure/backend software engineering.
  • •Experience designing and implementing large-scale distributed ML platform systems using Apache Spark, Apache Kafka, Apache Flink, or Databricks.
  • •Experience using modern programming languages including Python, Scala, Java, or Go for ML platform systems, backend services, data processing jobs, and automation tools.
  • •Experience with cloud platforms (AWS, Azure, or GCP) and infrastructure-as-code, containerization (Docker on managed orchestration like EKS/ECS), and monitoring using Prometheus and Grafana (including Grafana Mimir).
  • •Experience designing and building infrastructure for recommendation systems, moderation pipelines, or LLM serving/deployment systems, including Ray Serve or Triton; plus CI/CD and GitOps automation with tools such as Terraform, Terragrunt, Helm, Jenkins or Buildkite, and GitOps (including Scaffold).
Experience:5+ yearsMachine learningDistributed systemsPlatform engineeringData engineeringML infrastructureBackend engineeringCloudDevOps
Education:Bachelor's in Computer Science, Computer Engineering, or a related field
Skills:Technical guidanceMentorshipCross-functional collaborationTechnical documentationProblem-solving
Tech Stack:PythonScalaJavaGoApache SparkKafkaFlinkDatabricksAWSAzureGCPInfrastructure-as-codeDockerAmazon EKSAmazon ECSPrometheusGrafanaGrafana MimirRay ServeTriton

Company Brief

Match Group
Operates a portfolio of consumer dating products and services, including Tinder, Match.com, PlentyOfFish, and Hinge, connecting millions of users globally through mobile and web platforms.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Dallas, United States
Founded: 1995
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor