Applied Researcher, Vision Language Models/VLM - TikTok
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Problem-solving","Creative mindset","Cross-functional collaboration","Project execution"]Research and advance multimodal reasoning and generation for vision language models (VLMs) and Omni models. Work on VLM enhancements such as OCR and captioning, explore model architecture and inference-efficient design, and build scalable applications for TikTok business scenarios including content moderation, search, recommendations, and client AI. Collaborate cross-functionally to plan and implement research projects, with an emphasis on evaluations, data processing recipes, and reinforcement-learning-based alignment.

