为什么人们追逐无用的令牌保存插件,却忽视了真正的解决方案

2 分•作者: yohji1984•3 个月前
我昨天写了一篇博客,讨论了 RTK 和 Ponytail 在实际编码任务中的无用之处。并发布了我的代理(agent)在 80% 真实 token 节省情况下的长周期任务基准测试。 我想知道为什么人们会忽略这些插件无用的事实,并且不关心实际的节省效果? 完整报告请查看我的仓库:https://github.com/Tura-AI/tura | Arm n | Harness score | Total tokens | Modeled cost | Rounds | Duration | |---|---|---|---|---|---| | No plugin | 2 | 78.85% | 6.660M | $5.281946 | 62.5 | 895s | | Ponytail | 2 | 80.77% | -7.56% | -8.87% | -9.60% | +13.51% | | RTK | 2 | 76.92% | +13.20% | +7.18% | +44.00% | +40.69% | | Configuration | Passes | Pass rate | Observed tokens | Rounds | Estimated cost | |---|---|---|---|---|---| | Tura Balanced High | 48/60 | 80.0% | 229,695,477 | 2,017 | $221.138 | | Tura Direct High | 39/60 | 65.0% | 75,108,167 | 969 | $99.620 | | Codex CLI Medium | 38/60 | 63.3% | 333,538,349 | 3,140 | $257.173 | | Codex CLI High | 36/60 | 60.0% | 455,742,296 | 6,074 | $327.483 |
查看原文
I wrote a blog yesterday on how useless RTK and Ponytail are on real coding tasks. And published my agent harness long-horizon task benchmarks on 80% real token saving.<p>I just want to know why people just ignore the fact those pulgins are useless and don&#x27;t care about the real savings?<p>full reports are on my repo: https:&#x2F;&#x2F;github.com&#x2F;Tura-AI&#x2F;tura<p>Arm n Harness score Total tokens Modeled cost Rounds Duration No plugin 2 78.85% 6.660M $5.281946 62.5 895s Ponytail 2 80.77% -7.56% -8.87% -9.60% +13.51% RTK 2 76.92% +13.20% +7.18% +44.00% +40.69%<p>Configuration Passes Pass rate Observed tokens Rounds Estimated cost Tura Balanced High 48&#x2F;60 80.0% 229,695,477 2,017 $221.138 Tura Direct High 39&#x2F;60 65.0% 75,108,167 969 $99.620 Codex CLI Medium 38&#x2F;60 63.3% 333,538,349 3,140 $257.173 Codex CLI High 36&#x2F;60 60.0% 455,742,296 6,074 $327.483