Member of Technical Staff - Post Training, Applied (Vision)

Liquid AI
San Francisco, Boston, United States, Cambridge
Workplace: HybridFull timeFunction: Education & TrainingSkills: ["Communication","Problem-solving","Ownership","Teamwork"]

A hands-on ML engineering role owning end-to-end post-training for vision-language models (VLMs) in production. You’ll translate customer requirements into multimodal post-training specs, manage data generation and evaluation pipelines, and run fine-tuning/RL workflows, shaping how Liquid AI's multimodal stack deploys to enterprise clients.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Liquid AI
Liquid AI
5 months ago

Member of Technical Staff - Post Training, Applied (Vision)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

A hands-on ML engineering role owning end-to-end post-training for vision-language models (VLMs) in production. You’ll translate customer requirements into multimodal post-training specs, manage data generation and evaluation pipelines, and run fine-tuning/RL workflows, shaping how Liquid AI's multimodal stack deploys to enterprise clients.
Location: San Francisco, Boston, United States, Cambridge
Workplace: Hybrid
Employment Type: Full time · Permanent
Job Function: Education & Training

Key Responsibilities

  • •Act as the technical owner for enterprise customer VLM post-training engagements.
  • •Translate customer requirements into concrete multimodal post-training specifications and workflows.
  • •Design and execute visual data generation, filtering, and quality assessment processes, including image-text pair curation, annotation pipelines, and synthetic data generation for visual tasks.
  • •Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for vision-language models.
  • •Design task-specific evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities. Interpret results and feed learnings back into core post-training pipelines.

Pay and Benefits

Perks:Health Insurance401kPaid LeaveEquity

Key Requirements

  • •Hands-on experience with data generation and evaluation for VLM or multimodal post-training.
  • •Experience training or fine-tuning vision-language models using SFT, preference alignment, and/or RL.
  • •Strong intuition for visual data quality, annotation design, and multimodal evaluation.
  • •Familiarity with vision encoders, image-text architectures, and how visual representations interact with language model backbones.
  • •Experience with visual grounding, document understanding, OCR, or video understanding tasks.
Experience:MultimodalVision-language modelsArtificial intelligence
Skills:CommunicationProblem-solvingOwnershipTeamwork
Tech Stack:Vision-language modelsSFTReinforcement learningOCRImage encodersMultimodal pipelines

Company Brief

Liquid AI
Builds AI infrastructure and tooling to enable real-time, distributed machine learning and orchestration across edge and cloud environments, simplifying deployment and management of intelligent applications.
Industry: AI & Machine Learning
Website