跳到正文
原文
Hugging Face Blog·· 2024-10-29精选AI 评分65

Intel Labs 与 Hugging Face 推出 Universal Assisted Generation,跨模型家族加速解码 1.5x-2.0x

Universal Assisted Generation: Faster Decoding with Any Assistant Model

AI 导读

Intel Labs 与 Hugging Face 推出 Universal Assisted Generation(UAG),通过双向 tokenizer 翻译让辅助生成不再要求目标模型与助手模型共享同一 tokenizer,可将任意 decoder 或 MoE 模型的推理加速 1.5x-2.0x 且几乎零开销。

推荐理由

原文解释了跨 tokenizer 助手模型的实现原理并给出多组加速数据,还提供 Transformers 4.46.0 的可用代码。

来源:Hugging Face Blog · huggingface.co