Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: phdSkills: ["Collaboration","Communication","Performance optimization","System efficiency","Open-source innovation"]

Build and scale cloud-native infrastructure for large-scale LLM inference as part of the Inference Infrastructure team. You’ll design container-based cluster management and orchestration systems for extreme performance, resiliency, and cost efficiency, and architect GPU-optimized AI accelerator infrastructure. Collaborate to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and stay current with Kubernetes, Ray, and other open-source advances, contributing production-ready code and open-source efforts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and scale cloud-native infrastructure for large-scale LLM inference as part of the Inference Infrastructure team. You’ll design container-based cluster management and orchestration systems for extreme performance, resiliency, and cost efficiency, and architect GPU-optimized AI accelerator infrastructure. Collaborate to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and stay current with Kubernetes, Ray, and other open-source advances, contributing production-ready code and open-source efforts.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • •Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms.
  • •Collaborate across teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • •Track advances in open source (Kubernetes, Ray, etc.) and AI/ML infrastructure research, integrating best practices into production systems.
  • •Write production-ready code that is maintainable, testable, and scalable.

Key Requirements

  • •Completing or recently completing a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Strong understanding of large model inference, distributed/parallel systems, and/or high-performance networking.
  • •Hands-on experience building cloud or ML infrastructure (resource management, scheduling, request routing, monitoring, or orchestration).
  • •Solid knowledge of container and orchestration technologies such as Docker and Kubernetes.
  • •Proficiency in at least one major programming language: Go, Rust, Python, or C++.
Experience:LLM inferenceDistributed systemsCloud-nativeGPU accelerationKubernetesOpen source
Education:PhD / Doctorate
Skills:CollaborationCommunicationPerformance optimizationSystem efficiencyOpen-source innovation
Tech Stack:AIBrixKubernetesDockerRayVLLMSGLangTensorRT-LLMGoRustPythonC++CUDAAWSAzureGCPSageMakerAzure MLVertex AIDeepSpeedPyTorch

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn