Senior Cloud Acceleration Engineer – DPU & AI Infra

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Debugging","Performance optimization","Architecture design","Technical collaboration","Research-oriented problem solving"]

Build and optimize DPU network software for ByteDance’s Volcano Engine Public Cloud, targeting high performance, low latency, and reliability. Partner with hardware teams on software-hardware co-design for networking and storage acceleration, and explore AI/ML infrastructure acceleration for distributed training and inference. Own end-to-end performance improvements across OS kernels/drivers to user-space runtime systems, and contribute to architecture, technical proposals, and research roadmaps.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Cloud Acceleration Engineer – DPU & AI Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize DPU network software for ByteDance’s Volcano Engine Public Cloud, targeting high performance, low latency, and reliability. Partner with hardware teams on software-hardware co-design for networking and storage acceleration, and explore AI/ML infrastructure acceleration for distributed training and inference. Own end-to-end performance improvements across OS kernels/drivers to user-space runtime systems, and contribute to architecture, technical proposals, and research roadmaps.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design and develop DPU network software focused on high performance, low latency, and reliability.
  • •Collaborate with hardware teams on software-hardware co-design for networking and storage acceleration.
  • •Explore AI/ML infrastructure acceleration using DPUs, GPUs, and custom hardware for distributed training and inference.
  • •Drive end-to-end performance optimization across OS kernels/drivers to user-space runtime systems.
  • •Contribute to architecture design, technical proposals, and long-term research directions.

Key Requirements

  • •B.S./M.S. in Computer Science/Computer Engineering (or Ph.D. with strong research/publications).
  • •3+ years of relevant industry experience (exception for Ph.D. with strong background).
  • •Proficiency in C/C++ development and debugging.
  • •Strong Linux systems development and solid understanding of compute, network architecture, and operating systems.
  • •Background in software-hardware co-design, distributed systems, high-performance networking, or AI/ML systems.
Experience:3+ years
Education:
Skills:DebuggingPerformance optimizationArchitecture designTechnical collaborationResearch-oriented problem solving
Tech Stack:CC++LinuxDPDKRDMAOVSSR-IOVEBPFGPUCUDAFPGAASICNCCLHypervisorsGPU virtualizationDistributed storage accelerationAI/MLInferenceTrainingInference kv cache

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn