Staff Machine Learning Platform Engineer

Faire
Waterloo, Toronto, Ontario, San Francisco
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: mastersSkills: ["Technical documentation"]

Design, improve, and operate a scalable machine learning platform that accelerates model training, deployment, and governance. You’ll build ML infrastructure (workspaces, clusters, jobs, workflows), productionize workloads with Spark/Delta Lake/MLflow/Databricks, and implement data governance with Unity Catalog. Lead MLOps practices including CI/CD (Terraform, GitHub Actions), IAM/RBAC, observability for quality and performance, and maintain technical documentation for internal teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Faire
Faire
3 months ago

Staff Machine Learning Platform Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Design, improve, and operate a scalable machine learning platform that accelerates model training, deployment, and governance. You’ll build ML infrastructure (workspaces, clusters, jobs, workflows), productionize workloads with Spark/Delta Lake/MLflow/Databricks, and implement data governance with Unity Catalog. Lead MLOps practices including CI/CD (Terraform, GitHub Actions), IAM/RBAC, observability for quality and performance, and maintain technical documentation for internal teams.
Location: Waterloo, Toronto, Ontario, San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and operate ML infrastructure including workspaces, clusters, jobs, and workflows.
  • •Productionize ML workloads using Spark, Delta Lake, MLflow, and Databricks Workflows.
  • •Implement Unity Catalog for data governance, lineage, access control, and secure multi-tenant usage.
  • •Build CI/CD pipelines for ML using Terraform and Git-based workflows (e.g., GitHub Actions).
  • •Optimize performance, reliability, and cost across training and inference workloads while establishing observability for data quality, model performance, and platform health.

Pay and Benefits

Equity and Bonus:Equity

Key Requirements

  • •8+ years of experience building production ML or data platforms.
  • •Strong hands-on expertise with Databricks, Spark, Delta Lake, and MLflow.
  • •Proficiency in Python and SQL, plus distributed systems concepts.
  • •Experience with cloud platforms and infrastructure-as-code.
  • •Solid MLOps understanding across CI/CD, monitoring, reproducibility, and security.
Experience:8+ yearsMLOps
Education:Master's
Skills:Technical documentation
Languages:English
Tech Stack:PythonSQLKotlinPyTorchMLflowSparkDelta LakeDatabricks WorkflowsKafkaDatabricksSnowflakeFivetranIcebergUnity CatalogDatadogAirflowCockroach DBMySQLAWSS3

Company Brief

Faire
Operates a wholesale marketplace connecting independent retailers with emerging and established brands, providing ordering, logistics, and payment solutions to help small businesses discover and stock unique products.
Industry: Online Marketplaces
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Headquarters: San Francisco, United States
Founded: 2017
WebsiteLinkedIn