Hugging Face Blog·· 2025-04-16AI 评分53
TNG 解析并发请求下的 LLM Prefill 与 Decode 性能优化
Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
AI 导读
TNG 在 Hugging Face 博客发布 LLM 性能系列第二篇,讲解 prefill 与 decode 两个阶段对延迟、吞吐和 GPU 利用率的影响,并比较 static batching、continuous batching 和 prefill-first 等并发策略。
来源:Hugging Face Blog · huggingface.co