GPT-5.6 Sol Max 值得购买吗?

2 分•作者: yohji1984•3 个月前
我使用 gpt-5.6 sol 在 codex cli 和我自己的 agent harness 上进行了最大推理测试:https://github.com/Tura-AI/tura。我只测试了一个任务,token 效率的差异不像在高模式下那么大:在重写基准测试中,Tura 使用的轮次最多减少了 83.1%,并将 harness 的结果提高了 7 个百分点。在 ultra 模式下,差异约为 50%,并将 harness 的结果仅提高了 1%。 完整报告:https://turaai.net/blog#is-gpt-5-6-sol-max-worth-it 然而,在基于 DeepSWE 等发布的报告的 bug 修复任务中,与成本的增加相比,性能的提升并不那么令人印象深刻。 以下是总结: 构建新项目时使用 Max 模式,调试或添加已定义功能时使用 High 模式。
查看原文
I ran some test with gpt- 5.6 sol in max reasoning on both codex cli and my own agent harness: https:&#x2F;&#x2F;github.com&#x2F;Tura-AI&#x2F;tura I tested only 1 task and the toekn efficency difference is not as great as in high mode: Tura used up to 83.1% fewer turns on the rewrite benchmark and improved the harness results by 7 percent point. In ultra the dirrence is around 50% and incressed harness result only by one. Full report: https:&#x2F;&#x2F;turaai.net&#x2F;blog#is-gpt-5-6-sol-max-worth-it<p>However in the bug fixing tasks based on the pulbished report from DeepSWE and others, the performance incress is less impressive compare with the incess in costs. here is the summarry:<p>Use Max for building new project, and high for debuging or add defined feature.