Tech Lead Software Engineer - AI Compute Infrastructure
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Communication","Collaboration","Performance optimization","System efficiency"]Lead the design and development of large-scale, container-based cluster management and orchestration systems for LLM inference infrastructure. Architect cloud-native, GPU-accelerated platforms that are performant, scalable, resilient, secure, and cost-efficient. Collaborate across teams to integrate and operate inference solutions using vLLM, SGLang, and TensorRT-LLM, stay current with Kubernetes/Ray and ML systems research, and deliver production-ready, maintainable code.

