为什么人们追逐无用的令牌保存插件,却忽视了真正的解决方案
2 分•作者: yohji1984•3 个月前
我昨天写了一篇博客,讨论了 RTK 和 Ponytail 在实际编码任务中的无用之处。并发布了我的代理(agent)在 80% 真实 token 节省情况下的长周期任务基准测试。
我想知道为什么人们会忽略这些插件无用的事实,并且不关心实际的节省效果?
完整报告请查看我的仓库:https://github.com/Tura-AI/tura
| Arm n | Harness score | Total tokens | Modeled cost | Rounds | Duration |
|---|---|---|---|---|---|
| No plugin | 2 | 78.85% | 6.660M | $5.281946 | 62.5 | 895s |
| Ponytail | 2 | 80.77% | -7.56% | -8.87% | -9.60% | +13.51% |
| RTK | 2 | 76.92% | +13.20% | +7.18% | +44.00% | +40.69% |
| Configuration | Passes | Pass rate | Observed tokens | Rounds | Estimated cost |
|---|---|---|---|---|---|
| Tura Balanced High | 48/60 | 80.0% | 229,695,477 | 2,017 | $221.138 |
| Tura Direct High | 39/60 | 65.0% | 75,108,167 | 969 | $99.620 |
| Codex CLI Medium | 38/60 | 63.3% | 333,538,349 | 3,140 | $257.173 |
| Codex CLI High | 36/60 | 60.0% | 455,742,296 | 6,074 | $327.483 |
查看原文
I wrote a blog yesterday on how useless RTK and Ponytail are on real coding tasks. And published my agent harness long-horizon task benchmarks on 80% real token saving.<p>I just want to know why people just ignore the fact those pulgins are useless and don't care about the real savings?<p>full reports are on my repo: https://github.com/Tura-AI/tura<p>Arm n Harness score Total tokens Modeled cost Rounds Duration
No plugin 2 78.85% 6.660M $5.281946 62.5 895s
Ponytail 2 80.77% -7.56% -8.87% -9.60% +13.51%
RTK 2 76.92% +13.20% +7.18% +44.00% +40.69%<p>Configuration Passes Pass rate Observed tokens Rounds Estimated cost
Tura Balanced High 48/60 80.0% 229,695,477 2,017 $221.138
Tura Direct High 39/60 65.0% 75,108,167 969 $99.620
Codex CLI Medium 38/60 63.3% 333,538,349 3,140 $257.173
Codex CLI High 36/60 60.0% 455,742,296 6,074 $327.483