Software Engineer - AI Compute Infrastructure

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 2+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","System efficiency","Performance optimization","Scalability focus"]

Build and operate cloud-native infrastructure for large-scale LLM inference in ByteDance’s Inference Infrastructure team. Design high-performance, scalable container-based cluster management and GPU/AI accelerator orchestration systems, and integrate next-gen inference solutions using vLLM, SGLang, and TensorRT-LLM. Stay current with open source (Kubernetes, Ray) and AI/ML systems research, writing production-ready code for resilient, cost-efficient ML platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Software Engineer - AI Compute Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate cloud-native infrastructure for large-scale LLM inference in ByteDance’s Inference Infrastructure team. Design high-performance, scalable container-based cluster management and GPU/AI accelerator orchestration systems, and integrate next-gen inference solutions using vLLM, SGLang, and TensorRT-LLM. Stay current with open source (Kubernetes, Ray) and AI/ML systems research, writing production-ready code for resilient, cost-efficient ML platforms.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • •Architect cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms.
  • •Collaborate across teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • •Monitor and adopt advances in open source (Kubernetes, Ray), AI/ML/LLM infrastructure, and systems research into production.
  • •Write high-quality, production-ready code that is maintainable, testable, and scalable.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related fields with 2+ years of relevant experience (Ph.D. with strong systems/ML publications also considered).
  • •Strong understanding of large model inference, distributed/parallel systems, and/or high-performance networking systems.
  • •Hands-on experience building cloud or ML infrastructure (resource management, scheduling, request routing, monitoring, or orchestration).
  • •Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • •Proficiency in at least one major programming language (Go, Rust, Python, or C++).
Experience:2+ yearsLLM inferenceDistributed systemsCloud infrastructureMachine learningOpen source
Education:Bachelor's in Computer Science, Computer Engineering, or related fields
Skills:CommunicationCollaborationSystem efficiencyPerformance optimizationScalability focus
Tech Stack:KubernetesAIBrixDockerRayGoRustPythonC++VLLMSGLangTensorRT-LLMCUDAAWSAzureGCPSageMakerAzure MLVertex AIDeepSpeedPyTorch

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn