AI Frameworks Software Engineer – Model Compression

Intel
Shanghai
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: ["Teamwork","Collaboration","Problem-solving","Communication","Self-motivation"]

Join the Intel Neural Compressor team to develop model compression technologies, including AutoRound and core algorithm tools, and optimize them for Intel AI platforms across CPUs, GPUs, and AI accelerators. You will research and implement quantization and compression for LLMs and multimodal generative models, explore efficient deployment and inference acceleration, and contribute to fine-tuning acceleration and open-source initiatives.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Intel
Intel
5 hours ago

AI Frameworks Software Engineer – Model Compression

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Join the Intel Neural Compressor team to develop model compression technologies, including AutoRound and core algorithm tools, and optimize them for Intel AI platforms across CPUs, GPUs, and AI accelerators. You will research and implement quantization and compression for LLMs and multimodal generative models, explore efficient deployment and inference acceleration, and contribute to fine-tuning acceleration and open-source initiatives.
Location: Shanghai
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Develop Intel Neural Compressor and core algorithm tools, including AutoRound, and optimize them for Intel AI platforms.
  • •Research and implement quantization and compression techniques for LLMs, VLMs, and generative models.
  • •Work on efficient model deployment and inference acceleration, and explore fine-tuning acceleration directions.
  • •Collaborate across algorithm research, product feature development, and performance optimization.
  • •Contribute to the open-source community.

Key Requirements

  • •Bachelor’s or master’s degree in Computer Science or a related field.
  • •Solid understanding of deep learning and large language model (LLM) fundamentals.
  • •Familiarity with model compression techniques such as quantization and pruning.
  • •Proficiency in Python, C++, or other programming languages used in deep learning development.
  • •Strong teamwork mindset, collaboration skills, and good verbal/written English communication.
Experience:Model compressionLLMsOpen-source
Education:Bachelor's in Computer Science
Skills:TeamworkCollaborationProblem-solvingCommunicationSelf-motivation
Languages:English
Tech Stack:PythonC++Deep learningQuantizationPruningSparsityKnowledge distillationLow-precision trainingAutoRoundLLMsVLMText-to-imageText-to-video

Company Brief

Intel
Designs and manufactures semiconductor chips, processors, and related hardware for PCs, data centers, networking, and embedded applications, while providing software and services to accelerate computing across industries globally.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1968
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor