您是如何衡量 Claude Code 和 Codex 的性能的?

2 分•作者: achalpandey•3 个月前
我认为编码基准测试结果并不能反映我们混乱的现实。因为它们: 1) 使用专门构建的测试工具。我们使用 Claude Code 或 Codex。 2) 测试一次性任务。我们在会话中工作。 会话是混乱的,我们从一个大的主要任务开始,然后进行一些清理,在这里修复一个相邻的问题,在那里再修复一个。我们开始、停止,并且会改变主意。 这会改变成本和质量。缓存的 TTL 会过期。上下文会增长。 我正在着手创建一项。这是我的粗略计划: 1) 使用 Claude Code 和 Codex。 2) 使用会话形式的工作负载。将多个 SWE bench 验证过的任务缝合到一个大的会话中。 2.a) 使用来自同一存储库的任务,以确保主题的连续性。 3) 主要指标:美元成本与质量。 3.a) 次要指标:轮次计数,完成时间。 待解决的问题: a) 这个问题有意义吗? b) 我的基准测试规范有意义吗? c) 以一个困难的任务开始的 10 个 SWE bench 验证过的任务是正确的工作负载形状吗?
查看原文
I think coding benchmark results don&#x27;t represent our messy reality. As they<p>1) Use purpose-built test harnesses We use Claude Code or Codex<p>2) Test one-shot tasks We work in sessions<p>Sessions are messy, we start with a large primary task, then some cleanup, an adjacent fix here another over there. We start, stop, and change our minds.<p>That changes both cost and quality. Cache TTLs expire. Context grows.<p>I am working on creating one. Here&#x27;s my rough plan: 1) Use Claude Code and Codex 2) Use session shaped workloads. Stitch multiple SWE bench verified tasks into one big session. 2.a) Use tasks from the same repo to ensure topical continuity. 3) Headline metrics: dollar cost vs quality. 3.a) Secondary metrics: turn count, time to completion<p>Open Questions a) Does this problem make sense? b) Does my benchmark spec make sense? c) Are 10 SWE bench verified task starting with a hard one the right workload shape?