BigHugger
GH Repository · intel

intel-extension-for-transformers

⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡

stars
2,172
30-day movement
starts with the next reading
Related entries
60
Connections
1
Pythongaudi3autoroundchatbotlarge-language-model4-bitspythonspeculative-decodingneural-chat-7bllm-inferencellm-cpuchatpdfretrievalragstreamingllmhabanaintel-optimized-llamacppneural-chat

Intel Extension for Transformers is a Python toolchain for building chatbots and running LLM inference efficiently on Intel platforms, including CPUs and Habana/Gaudi accelerators. It provides state-of-the-art model compression techniques such as 4-bit quantization (AutoRound), speculative decoding, and StreamingLLM support, plus RAG capabilities like ChatPDF.

Use it when you want to deploy or compress LLMs and chatbots specifically optimized for Intel hardware.

Use it to

  • Build a chatbot on Intel devices
  • Quantize LLMs to 4-bit with AutoRound
  • Run ChatPDF-style RAG pipelines
  • Serve LLM inference on Intel CPUs
  • Deploy on Habana Gaudi accelerators

For Developers deploying LLMs on Intel hardware

Role
rag
Language
Python
Licence
Apache-2.0
Forks
215
Open issues
31
Last push
2024-10-08
Latest release
v1.0a · 2022-11-23
topicsllm-inferencequantizationragchatbotintelcompression