Member of Technical Staff - Multi-Modal, Vision
San Francisco
Workplace: RemoteFull timeFunction: Software EngineeringEducation: mastersSkills: ["Communication","Problem-solving","Teamwork"]Join the VLM team to own end-to-end development of vision-language models that run on-device, balancing latency/memory constraints with model quality. You’ll drive data curation, training recipes, ablations, and evaluation, collaborating with pretraining/post-training and infrastructure teams to ship state-of-the-art models. Opportunities include leading a major work-stream and pushing token-efficient encoders.

