我已对 llama.cpp 进行补丁,将提示处理 TPS 提升了 20%。请帮我创建一个 PR。
4 分•作者: i_am_rocoe•4 个月前
我在本地使用 llama.cpp 运行 Qwen3.6-35B-A3B,注意到在使用 MTP 时,提示处理吞吐量会变得非常低。我对此产生了浓厚的兴趣。
最初的好奇心变成了一个为期两周的实验深渊,最终我完成了一个概念验证(PoC),该 PoC 完全消除了 GPU 上 MTP 的提示处理(PP)开销,超出了我所有的预期。
简而言之:这个 PoC 没有像往常一样处理整个批次(ubatch)的 token(通常是 512-2048 个 token)的最后一层 MoE FFN,而是只处理输出行(预填充期间通常是 1 个 token)。结果是,提示处理(PP)的每秒吞吐量(TPS)恢复到禁用 MTP 时的水平,同时保留了 MTP 对总吞吐量(TG)的大部分好处,即使在其中一个基准测试中草稿接受率略有下降。
我不会向 llama.cpp 提交拉取请求(PR),因为这是 AI 生成的代码,这违反了他们的贡献政策,我对此表示支持。如果你懂 C++ 和 llama.cpp 的内部原理,我邀请你与我合作,共同提交一个更成熟的实现版本。
查看原文
I've been running Qwen3.6-35B-A3B locally on llama.cpp and noticed that prompt processing throughput gets too low with MTP. I got nerd-sniped.<p>What started as curiosity turned into a two-week rabbit hole of experiments and ended with a PoC that fully recovers the MTP PP overhead on GPU, above any expectation I had.<p>TL;DR: instead of processing the last layer MoE FFN for the entire ubatch tokens (usually 512-2048 tokens), this PoC processes only the output row (usually 1 token during prefill). The result is PP TPS is back to the same as with MTP disabled, keeping most of MTP's benefits to TG TPS, even with a slight drop in draft acceptance rate in one of the benchs.<p>I'm not opening a PR to llama.cpp because this is AI-generated code, which goes against their contribution policy, which I support. If you know C++ and llama.cpp internals, I invite to work together with me to open a PR with a more mature implementation.