跳到正文
原文
Hugging Face Blog·· 2022-08-17精选AI 评分74

Hugging Face 详解 LLM.int8():用 bitsandbytes 实现 8-bit 量化推理

A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes

AI 导读

Hugging Face 与 BigScience 将 LLM.int8() 8-bit 量化集成进 transformers 和 accelerate,使 BLOOM-176B 等 176B 参数大模型的推理显存占用减半且性能不下降。

推荐理由

原文由集成团队讲解 LLM.int8() 的量化原理、零性能下降验证和 transformers 集成细节,读者可据此在有限显存上运行大模型。

来源:Hugging Face Blog · huggingface.co