跳到正文
原文
Hugging Face Blog·· 2026-08-10AI 评分43

Hugging Face 让知识蒸馏成本低到可大规模运行

Making Knowledge Distillation Cheap Enough to Run at Scale

AI 导读

Hugging Face 论文提出两项系统改动来降低 LLM 知识蒸馏成本:离线缓存 teacher 的 top-100 logits,以及将输出投影融合进损失的 fused chunked KL loss,避免构建全词表 × 序列长度矩阵。

来源:Hugging Face Blog · huggingface.co