Senior Software Engineer, Machine Learning Systems (Multiple Positions)

ByteDance
San Jose
Workplace: OnsiteFull timeUSD 212,800 - 387,600 annuallyFunction: Software EngineeringExperience: 2-5 yearsEducation: mastersSkills: ["Curiosity","Humility"]

Design, develop, and optimize machine learning systems with a focus on heterogeneous computing, resources management, and system monitoring. Build and maintain scalable ML infrastructure including distributed task scheduling and large-scale training pipelines. Drive cross-layer optimization across hardware, systems, and AI algorithms to improve performance and efficiency. Enhance training frameworks for general-purpose and model-specific needs such as LLMs and diffusion models, improving reliability and throughput at massive scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Software Engineer, Machine Learning Systems (Multiple Positions)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Design, develop, and optimize machine learning systems with a focus on heterogeneous computing, resources management, and system monitoring. Build and maintain scalable ML infrastructure including distributed task scheduling and large-scale training pipelines. Drive cross-layer optimization across hardware, systems, and AI algorithms to improve performance and efficiency. Enhance training frameworks for general-purpose and model-specific needs such as LLMs and diffusion models, improving reliability and throughput at massive scale.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, develop, and optimize machine learning systems across heterogeneous computing architectures.
  • •Deploy and maintain scalable ML infrastructure, including distributed task scheduling and large-scale training pipelines.
  • •Facilitate cross-layer optimization across hardware, systems, and AI algorithms to improve ML workload performance and efficiency.
  • •Implement and enhance training framework features for both general-purpose and model-specific optimizations (e.g., LLMs and diffusion models).
  • •Improve reliability, efficiency, and throughput for massive-scale distributed training jobs, and mentor interns/junior engineers.

Pay and Benefits

Salary: USD 212,800 - 387,600 annually

Key Requirements

  • •Master’s degree (or foreign equivalent) in a relevant field with 2 years of related experience, or a Bachelor’s degree (or foreign equivalent) with 5 years of related experience.
  • •Design and implement software service architecture for high-throughput and fault-tolerant services.
  • •Work with GPU or ASIC-based computing systems for machine learning applications.
  • •Develop ML models using platforms/frameworks including PyTorch or JAX.
  • •Program in Linux using C/C++, CUDA, or Python, and design/manage databases for online services using MySQL or Redis.
Experience:2-5 yearsMachine learningDistributed systemsDeep learningLarge-scale training
Education:Master's in Computer Science, Engineering (any), Information Systems, Data Science, Mathematics, or a related field
Skills:CuriosityHumility
Tech Stack:PyTorchJAXMySQLRedisLinuxC/C++CUDAGPUASICDistributed task schedulingTraining pipelinesLarge language modelsDiffusion models

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn