GH Repository · intel
intel-extension-for-transformers
⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡
- stars
- 2,172
- 30-day movement
- starts with the next reading
- Related entries
- 60
- Connections
- 1
Pythongaudi3autoroundchatbotlarge-language-model4-bitspythonspeculative-decodingneural-chat-7bllm-inferencellm-cpuchatpdfretrievalragstreamingllmhabanaintel-optimized-llamacppneural-chat
Intel Extension for Transformers is a Python toolchain for building chatbots and running LLM inference efficiently on Intel platforms, including CPUs and Habana/Gaudi accelerators. It provides state-of-the-art model compression techniques such as 4-bit quantization (AutoRound), speculative decoding, and StreamingLLM support, plus RAG capabilities like ChatPDF.
Use it when you want to deploy or compress LLMs and chatbots specifically optimized for Intel hardware.
Use it to
- Build a chatbot on Intel devices
- Quantize LLMs to 4-bit with AutoRound
- Run ChatPDF-style RAG pipelines
- Serve LLM inference on Intel CPUs
- Deploy on Habana Gaudi accelerators
For Developers deploying LLMs on Intel hardware
- Role
- rag
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 215
- Open issues
- 31
- Last push
- 2024-10-08
- Latest release
- v1.0a · 2022-11-23
topicsllm-inferencequantizationragchatbotintelcompression