Senior Research Scientist - Machine Learning System

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Collaboration","Decision-making","Performance analysis"]

Build and optimize large-scale LLM inference frameworks and high-performance inference engines. Work on a distributed ML system that integrates heterogeneous compute (GPU/NPU), low-level performance (CUDA), and scalable infrastructure such as RDMA and storage. Collaborate with a global team across the United States, China, and Singapore to keep training/inference services stable, reliable, and scalable while improving coding, performance analysis, and distributed system decisions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Research Scientist - Machine Learning System

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize large-scale LLM inference frameworks and high-performance inference engines. Work on a distributed ML system that integrates heterogeneous compute (GPU/NPU), low-level performance (CUDA), and scalable infrastructure such as RDMA and storage. Collaborate with a global team across the United States, China, and Singapore to keep training/inference services stable, reliable, and scalable while improving coding, performance analysis, and distributed system decisions.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Develop and optimize LLM inference frameworks.
  • •Optimize GPU and CUDA performance to build a high-performance LLM inference engine.
  • •Build and maintain massively distributed ML training and inference system/services.
  • •Integrate large-scale heterogeneous systems using GPU/NPU/RDMA/Storage and ensure stability and reliability.
  • •Contribute to unified project direction within a global team working on ML systems.

Key Requirements

  • •Bachelor’s degree or above in computer/electronics/automation/software or related fields.
  • •Proficient in C/C++, algorithms, data structures, and familiar with Python.
  • •Understand deep learning algorithms, neural network architectures, and training frameworks such as PyTorch.
  • •Proficient in GPU high-performance computing optimization, including CUDA and parallel/memory access optimization (preferred).
  • •Familiar with LLM acceleration/optimization tools and model optimization, including TensorRT-LLM, ORCA, or VLLM (preferred).
Experience:Machine learningLLMDistributed systemsGPU computingLLM inference
Education:Bachelor's in computer/electronics/automation/software (or related)
Skills:CollaborationDecision-makingPerformance analysis
Tech Stack:C/C++PythonCUDAGPUNPURDMAStoragePyTorchTensorRT-LLMORCAVLLMNeural networksDeep learningLLM inference framework

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn