Staff Engineer - ML Infra / MLOps

Quince
Palo Alto
Workplace: OnsiteFull timeUSD 218,000 - 285,000 annuallyFunction: Software EngineeringSkills: ["Mentorship","Technical leadership","On-call discipline","Root-cause analysis","Operational excellence"]

Build and operate Quince’s ML infrastructure foundation, covering model training, feature pipelines, and high-throughput inference serving. Own the “paved road” experience that helps data scientists and AI researchers move from idea to production reliably. Set CI/CD and deployment standards (IaC, model versioning, experiment tracking), optimize GPU utilization and cloud costs, and ensure scalability through monitoring, alerting, and automated recovery.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Quince
Quince
1 month ago

Staff Engineer - ML Infra / MLOps

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and operate Quince’s ML infrastructure foundation, covering model training, feature pipelines, and high-throughput inference serving. Own the “paved road” experience that helps data scientists and AI researchers move from idea to production reliably. Set CI/CD and deployment standards (IaC, model versioning, experiment tracking), optimize GPU utilization and cloud costs, and ensure scalability through monitoring, alerting, and automated recovery.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own end-to-end design of Quince’s ML platform, including model training, serving, feature pipelines, and monitoring.
  • •Build the core developer experience that helps data scientists and AI researchers move from idea to production with minimal friction.
  • •Set and uphold ML engineering standards across CI/CD, Infrastructure as Code, model versioning, experiment tracking, and deployment strategies (blue-green, canary).
  • •Lead evaluation and selection of core platform components, balancing build vs. buy decisions (e.g., inference runtimes, feature stores, orchestration frameworks).
  • •Optimize GPU utilization and cloud costs, ensure production scalability and reliability via monitoring/alerting/recovery, and mentor engineers through reviews and pairing.

Pay and Benefits

Salary: USD 218,000 - 285,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •8+ years of industry experience with at least 4+ years hands-on in ML Infrastructure, MLOps, or large-scale Data Platform engineering.
  • •Proven ability to design and build MLOps platforms across the full model lifecycle, from ingestion and distributed training to real-time inference and governance.
  • •Deep expertise in cloud-native infrastructure (preferably AWS), Kubernetes (EKS), Docker, and Infrastructure as Code (Terraform/Pulumi).
  • •Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker.
  • •Expertise in ML CI/CD, including model versioning, experiment tracking, and deployment strategies like blue-green and canary.
Experience:ML infrastructureMLOpsData platforms
Skills:MentorshipTechnical leadershipOn-call disciplineRoot-cause analysisOperational excellence
Languages:English
Tech Stack:AWSKubernetesEKSDockerTerraformPulumiPyTorchTensorFlowKubeflowSageMakerSparkFlinkKafkaCI/CDInfrastructure as CodeFeature storesBlue-greenCanary

Company Brief

Quince
Quince is a direct-to-consumer brand offering affordable, high-quality apparel, home goods, and essentials—such as cashmere, basics, and home textiles—focused on transparent sourcing and value pricing for customers.
Industry: Direct to Consumer Brands
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: San Francisco, United States
Founded: 2015
WebsiteLinkedIn