Senior Software Engineer - AI Compute Infrastructure

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: mastersSkills: ["Communication","Collaboration","Performance optimization"]

Build and operate cloud-native, GPU-optimized infrastructure for large-scale LLM inference. You’ll design container-based cluster management and orchestration systems, architect secure and cost-efficient GPU/AI accelerator platforms, and collaborate with teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other engines. Work in a hyper-scale environment, keep up with open-source and systems research, and contribute production-ready, maintainable code for globally distributed datacenters.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Software Engineer - AI Compute Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate cloud-native, GPU-optimized infrastructure for large-scale LLM inference. You’ll design container-based cluster management and orchestration systems, architect secure and cost-efficient GPU/AI accelerator platforms, and collaborate with teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other engines. Work in a hyper-scale environment, keep up with open-source and systems research, and contribute production-ready, maintainable code for globally distributed datacenters.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • •Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms.
  • •Collaborate across teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • •Stay current with open source (Kubernetes, Ray), AI/ML, LLM infrastructure, and systems research; integrate best practices into production systems.
  • •Write production-ready code that is maintainable, testable, and scalable.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related fields with 3+ years of relevant experience (Ph.D. with strong systems/ML publications also considered).
  • •Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems.
  • •Hands-on experience building cloud or ML infrastructure across areas like resource management, scheduling, request routing, monitoring, or orchestration.
  • •Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • •Proficiency in at least one major programming language (Go, Rust, Python, or C++).
Experience:3+ yearsLLM inferenceDistributed systemsCloud infrastructureKubernetesGPU acceleration
Education:Master's in Computer Science, Computer Engineering, or related fields
Skills:CommunicationCollaborationPerformance optimization
Tech Stack:KubernetesAIBrixDockerRayVLLMSGLangTensorRT-LLMGoRustPythonC++CUDAAWSAzureGCPSageMakerAzure MLVertex AIDeepSpeedPyTorch

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn