Software Engineer — Distributed LLM Inference Systems

Intel
Shanghai
Workplace: OnsiteFull timeFunction: Software Engineering0Education: mastersSkills: ["Problem-solving","Performance optimization","Debugging","Collaboration","Communication"]

Design, develop, and optimize distributed inference systems for large language models on Intel’s AI Frameworks team. You’ll implement distributed inference algorithms, optimize model execution and communication, and build components such as request schedulers, KV cache management, and communication layers. Profile workloads to address latency, throughput, scalability, and resource utilization, and contribute production-quality code, tests, and documentation to internal and open-source projects.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Intel
Intel
17 hours ago

Software Engineer — Distributed LLM Inference Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Design, develop, and optimize distributed inference systems for large language models on Intel’s AI Frameworks team. You’ll implement distributed inference algorithms, optimize model execution and communication, and build components such as request schedulers, KV cache management, and communication layers. Profile workloads to address latency, throughput, scalability, and resource utilization, and contribute production-quality code, tests, and documentation to internal and open-source projects.
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Entry level

Key Responsibilities

  • •Design, develop, and optimize distributed LLM inference systems and AI framework components.
  • •Implement distributed algorithms, including model/data parallel frameworks and asynchronous communication for deep learning.
  • •Develop and optimize components such as request schedulers, model workers, communication layers, and KV cache management mechanisms.
  • •Profile distributed inference workloads to identify computation, communication, memory, and scheduling bottlenecks.
  • •Collaborate with component teams to improve end-to-end latency, throughput, scalability, and resource utilization, and contribute code, tests, and documentation to internal and open-source projects.

Key Requirements

  • •Master’s degree in computer science, Artificial Intelligence, Software Engineering, or related field with 0–1 years hands-on experience via internships, projects, coursework, or training.
  • •Proficiency in Python and modern C++ programming.
  • •Foundational knowledge of deep learning and AI frameworks such as PyTorch.
  • •Experience debugging and optimizing software for performance.
  • •Basic understanding of machine learning algorithms and strong problem-solving with the ability to learn unfamiliar systems quickly.
Experience:0Deep learningMachine learningLLM inferenceDistributed systemsOpen source
Education:Master's
Skills:Problem-solvingPerformance optimizationDebuggingCollaborationCommunication
Languages:English
Tech Stack:PythonC++PyTorchVLLMSGLangTensorRT-LLMLLM inferenceKV cache managementContinuous batchingParallelismDisaggregated servingRequest schedulingAsynchronous communicationDeep learning

Company Brief

Intel
Designs and manufactures semiconductor chips, processors, and related hardware for PCs, data centers, networking, and embedded applications, while providing software and services to accelerate computing across industries globally.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1968
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor