Member of Technical Staff - VLM

Black Forest Labs
Freiburg
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Communication","Teamwork","Problem-solving","Research"]

Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack, innovating on architectures rather than applying existing ones. You’ll design fine-tuning strategies for creative use cases, research integrations with diffusion/flow pipelines to boost generation quality at scale, and evaluate emerging architectures to translate research into practical improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Black Forest Labs
Black Forest Labs
3 months ago

Member of Technical Staff - VLM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack, innovating on architectures rather than applying existing ones. You’ll design fine-tuning strategies for creative use cases, research integrations with diffusion/flow pipelines to boost generation quality at scale, and evaluate emerging architectures to translate research into practical improvements.
Location: Freiburg
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack — innovating on architectures, not just applying existing ones
  • •Design fine-tuning strategies that adapt VLMs to specialized creative use cases (captioning, editing instructions, prompt enhancement) that general-purpose models can’t handle
  • •Research integrations between VLM/LLM capabilities and our diffusion and flow pipelines — finding creative ways to improve generation quality and controllability without computational bottlenecks
  • •Evaluate emerging multimodal architectures, translating the best of recent research into practical improvements
  • •Collaborate with distributed teams to ensure research translates into scalable production deployment

Key Requirements

  • •You have pretrained or significantly advanced a VLM (not just SFT’d or LoRA’d one) deployed in production or released publicly
  • •Strong publication record or unambiguous production track record showing you push the frontier on multimodal architectures
  • •Deep understanding of how vision and language representations interact: tokenization, alignment, grounding, cross-modal attention, and failure modes
  • •Experience with distributed training at multi-node scale
  • •Comfortable at the research/production boundary—you care whether the work ships and generalizes, not just whether it reads well
Experience:MultimodalGenerative AI
Skills:CommunicationTeamworkProblem-solvingResearch
Languages:English
Tech Stack:Vision-Language ModelsDiffusionFlow-based modelsDistributed trainingLLMs

Company Brief

Black Forest Labs
Develops AI-driven solutions and provides machine learning consultancy to help businesses integrate advanced models and data-driven workflows. Focuses on prototyping, model deployment, and bespoke AI services to address industry-specific problems.
Industry: Consulting
Website