Machine Learning Engineer, AML - Engine (Multiple Positions)

ByteDance
San Jose
Workplace: OnsiteFull timeUSD 187,741 - 316,800 annuallyFunction: Data Science & Machine LearningExperience: 1-3 yearsEducation: bachelorsSkills: []

Build and optimize large-scale machine learning systems for deep learning model architectures, applying LLM and reinforcement learning research to improve accuracy and efficiency. Design and implement high-performance CUDA kernels and parallel computing solutions to reduce latency and improve throughput. Analyze computation graphs for optimization opportunities and migrate model serving between TensorFlow and PyTorch while maintaining correctness and performance. Balance AUC with execution efficiency across in-graph and out-of-graph computations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
12 hours ago

Machine Learning Engineer, AML - Engine (Multiple Positions)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Build and optimize large-scale machine learning systems for deep learning model architectures, applying LLM and reinforcement learning research to improve accuracy and efficiency. Design and implement high-performance CUDA kernels and parallel computing solutions to reduce latency and improve throughput. Analyze computation graphs for optimization opportunities and migrate model serving between TensorFlow and PyTorch while maintaining correctness and performance. Balance AUC with execution efficiency across in-graph and out-of-graph computations.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Apply machine learning and robotics research knowledge, including LLM and reinforcement learning, to analyze and optimize deep learning model architectures.
  • •Design and implement high-performance CUDA kernels to accelerate model inference and computation-heavy operators.
  • •Analyze computation graphs to identify optimization opportunities across different model structures.
  • •Use parallel computing techniques such as CUDA and OpenMP to improve throughput, latency, and resource utilization in ML workloads.
  • •Migrate between TensorFlow and PyTorch model serving frameworks and develop algorithm-level optimizations balancing AUC and execution efficiency.

Pay and Benefits

Salary: USD 187,741 - 316,800 annually

Key Requirements

  • •Master's degree (or foreign equivalent) in Computer Science/Engineering/IT/Data Science/Robotics/Mathematics (or related quantitative field) with 1 year of related experience, or Bachelor's degree (or foreign equivalent) with 3 years of related experience.
  • •Python or C++ coding experience.
  • •Optimize deep learning training and inference pipelines on GPU and CPU architectures.
  • •Implement CUDA-based performance optimizations for machine learning workloads.
  • •Develop and deploy large-scale machine learning systems and use big data frameworks including MySQL, Spark, Hadoop, or Flink.
Experience:1-3 yearsMachine learningDeep learningLLMsReinforcement learningBig data
Education:Bachelor's in Computer Science, Engineering (any), Information Technology, Data Science, Robotics, Mathematics, or a related quantitative field
Tech Stack:PythonC++CUDATensorFlowPyTorchOpenMPMySQLSparkHadoopFlinkGPUCPU

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn