Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Communication","Collaboration","Performance optimization","System efficiency"]

Lead the design and development of large-scale, container-based cluster management and orchestration systems for LLM inference infrastructure. Architect cloud-native, GPU-accelerated platforms that are performant, scalable, resilient, secure, and cost-efficient. Collaborate across teams to integrate and operate inference solutions using vLLM, SGLang, and TensorRT-LLM, stay current with Kubernetes/Ray and ML systems research, and deliver production-ready, maintainable code.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Tech Lead Software Engineer - AI Compute Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Lead the design and development of large-scale, container-based cluster management and orchestration systems for LLM inference infrastructure. Architect cloud-native, GPU-accelerated platforms that are performant, scalable, resilient, secure, and cost-efficient. Collaborate across teams to integrate and operate inference solutions using vLLM, SGLang, and TensorRT-LLM, stay current with Kubernetes/Ray and ML systems research, and deliver production-ready, maintainable code.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • •Architect cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms.
  • •Collaborate across teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • •Monitor and integrate advances in open source and LLM infrastructure research (e.g., Kubernetes, Ray) into production systems.
  • •Write production-ready, maintainable, testable, and scalable code.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related fields with 5+ years of relevant experience (Ph.D. with strong systems/ML publications considered).
  • •Strong understanding of large model inference and distributed/parallel systems, including high-performance networking systems.
  • •Hands-on experience building cloud or ML infrastructure across resource management, scheduling, request routing, monitoring, or orchestration.
  • •Solid knowledge of container and orchestration technologies, including Docker and Kubernetes.
  • •Proficiency in at least one major programming language: Go, Rust, Python, or C++.
Experience:LLM inferenceDistributed systemsCloud infrastructureKubernetesGPU acceleration
Skills:CommunicationCollaborationPerformance optimizationSystem efficiency
Tech Stack:KubernetesRayAIBrixDockerVLLMSGLangTensorRT-LLMCUDAAWSAzureGCPSageMakerAzure MLVertex AIDeepSpeedPyTorchGoRustPythonC++

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn