Edge ML Software Engineer (Model Optimization-PICO) - San Jose

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: mastersSkills: []

Build and optimize edge AI pipelines for AR/VR imaging by converting and compiling ML models for execution on edge NPUs. Profile and analyze model performance and power across simulators, emulators, and silicon, then identify compute, memory, data movement, and scheduling bottlenecks. Apply hardware-aware optimizations like quantization, compression, and operator fusion to meet latency, memory, and power targets while partnering with algorithm, compiler, firmware, and hardware teams to debug issues.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Edge ML Software Engineer (Model Optimization-PICO) - San Jose

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize edge AI pipelines for AR/VR imaging by converting and compiling ML models for execution on edge NPUs. Profile and analyze model performance and power across simulators, emulators, and silicon, then identify compute, memory, data movement, and scheduling bottlenecks. Apply hardware-aware optimizations like quantization, compression, and operator fusion to meet latency, memory, and power targets while partnering with algorithm, compiler, firmware, and hardware teams to debug issues.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Convert and compile ML models for execution on edge NPUs and apply quantization mechanisms.
  • •Profile and analyze model performance and power consumption on simulators, emulators, and silicon platforms.
  • •Identify bottlenecks related to compute, memory bandwidth, data movement, and scheduling.
  • •Apply hardware-aware optimization strategies (quantization, compression, operator fusion) to meet latency, memory, and power targets.
  • •Collaborate with algorithm, compiler, firmware, and hardware teams to debug functional and performance issues.

Key Requirements

  • •Master's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.
  • •3+ years of industry experience in machine learning software engineering, model deployment, or ML systems for production environments.
  • •Strong understanding of deep learning architectures, including CNNs and Transformers.
  • •Knowledge of ML accelerators' architectures, operator fusion, memory hierarchies, and data movements.
  • •Practical experience with popular ML frameworks such as PyTorch or TensorFlow, plus proficiency in Python and C/C++.
Experience:3+ yearsMachine learningML systemsModel deploymentProduction environments
Education:Master's
Tech Stack:PythonC/C++PyTorchTensorFlowCNNsTransformersQuantizationOperator fusionPTQQATEmulatorsSimulators

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn