Hugging Face Blog·· 2022-08-17精选AI 评分74
Hugging Face 详解 LLM.int8():用 bitsandbytes 实现 8-bit 量化推理
A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
AI 导读
Hugging Face 与 BigScience 将 LLM.int8() 8-bit 量化集成进 transformers 和 accelerate,使 BLOOM-176B 等 176B 参数大模型的推理显存占用减半且性能不下降。
推荐理由
原文由集成团队讲解 LLM.int8() 的量化原理、零性能下降验证和 transformers 集成细节,读者可据此在有限显存上运行大模型。
来源:Hugging Face Blog · huggingface.co