Senior Machine Learning Operations Engineer

Mercury
San Francisco, New York
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Product ownership","Cross-functional collaboration","Operational excellence","Experimentation mindset","Attention to observability"]

Own the production lifecycle for machine learning models that power fraud and financial-crime risk decisions. Build and operate low-latency, high-availability real-time inference for the risk decision engine, including deployment infrastructure, CI/CD, and staged rollouts. Drive model observability with latency/error monitoring and drift detection, partner with risk data science for development-to-production handoffs, and implement experimentation and explainability capabilities like SHAP.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mercury
Mercury
5 days ago

Senior Machine Learning Operations Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Own the production lifecycle for machine learning models that power fraud and financial-crime risk decisions. Build and operate low-latency, high-availability real-time inference for the risk decision engine, including deployment infrastructure, CI/CD, and staged rollouts. Drive model observability with latency/error monitoring and drift detection, partner with risk data science for development-to-production handoffs, and implement experimentation and explainability capabilities like SHAP.
Location: San Francisco, New York
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Build and operate the real-time inference service that scores models for the risk decision engine with low latency and high availability.
  • •Own model deployment infrastructure including registry/versioning, CI/CD with performance, bias, and consistency checks, shadow mode, and staged rollouts.
  • •Implement model observability with availability, latency, and error monitoring, plus drift detection to trigger retraining.
  • •Partner with Risk Data Science to operate the development-to-production handoff and manage production operations under ML Platform ownership.
  • •Deliver experimentation and explainability capabilities such as champion/challenger, canary routing, and SHAP attributions.

Pay and Benefits

Equity and Bonus:Equity

Key Requirements

  • •5+ years in machine learning engineering, backend software engineering, MLOps, or a closely related field.
  • •Production ML service experience deploying, serving, and operating models in low-latency, high-availability contexts.
  • •Strong backend engineering fundamentals in Python, including API frameworks such as FastAPI or Flask.
  • •Experience with model deployment tooling and lifecycle patterns (registries, CI/CD, versioning, shadow/canary/champion-challenger, staged rollouts).
  • •Experience building observability and alerting for production services, including latency/errors and model drift signals.
Skills:Product ownershipCross-functional collaborationOperational excellenceExperimentation mindsetAttention to observability
Tech Stack:PythonFastAPIFlaskSHAPSQLRedisDynamoDBKafkaKinesisRedpandaSnowflakeDbtDagsterAirflowHaskellReactTypeScript

Company Brief

Mercury
Provides banking and financial infrastructure for startups and small businesses, offering business checking accounts, payments, debit cards, venture debt integrations, and treasury services via a developer-friendly platform.
Industry: Fintech Infrastructure
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2017
WebsiteLinkedIn