Research Scientist Intern - Multimodal Sensing & On-Device Perception - Global Frontier Tech Recruitment Program - 2027 Start (PHD)

ByteDance
San Jose
Workplace: OnsiteInternshipFunction: Data Science & Machine LearningEducation: phdSkills: ["Research","Problem-solving","Curiosity","Self-motivated"]

Work on the eye tracking system architecture stack, co-designing intelligent sensing hardware and perception algorithms for high coverage, performance, and low power. You’ll prototype novel sensor/imaging architectures, build end-to-end imaging pipelines from optical/sensor physics through ISP to downstream models, and use VLM/LLM and world model requirements to guide what information the sensing front-end preserves or transforms. Collaborate on machine vision models co-optimized for power, bandwidth, and latency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Intern - Multimodal Sensing & On-Device Perception - Global Frontier Tech Recruitment Program - 2027 Start (PHD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Work on the eye tracking system architecture stack, co-designing intelligent sensing hardware and perception algorithms for high coverage, performance, and low power. You’ll prototype novel sensor/imaging architectures, build end-to-end imaging pipelines from optical/sensor physics through ISP to downstream models, and use VLM/LLM and world model requirements to guide what information the sensing front-end preserves or transforms. Collaborate on machine vision models co-optimized for power, bandwidth, and latency.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Design and prototype novel sensor or imaging architectures that move computation closer to the sensing front-end (e.g., near-sensor processing, event-driven capture, learned pixel-level compression).
  • •Build and characterize end-to-end imaging pipelines from optical/sensor physics through ISP to downstream perception models, identifying where data/compute efficiency can improve.
  • •Use knowledge of VLM/LLM and world models to determine what sensing information should be preserved, discarded, or transformed.
  • •Develop or adapt machine vision models co-optimized with hardware constraints such as power, bandwidth, and latency.

Key Requirements

  • •Currently pursuing a PhD in Computer Science, Electrical Engineering, Optical Engineering, Applied Mathematics, Physics, or a related technical field.
  • •Strong research background in computer vision and machine learning, with hands-on model training experience.
  • •Experience with at least one of: sequence modeling, language modeling, efficient neural network design, or signal processing.
  • •Proven track record of high-impact research with publications in top venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, or SIGGRAPH.
  • •Hands-on experience prototyping hardware with cameras, structured light, or other active sensing systems.
Education:PhD / Doctorate in PhD (Computer Science, Electrical Engineering, Optical Engineering, Applied Mathematics, Physics, or related)
Skills:ResearchProblem-solvingCuriositySelf-motivated

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn