HN 提问:有人有兴趣构建一个仅限 Harness 的基准测试吗?

1作者: GodelNumbering大约 15 小时前
尽管有许多大型语言模型(LLM)的基准测试,但很少有(甚至没有)利用“harness”进行基准测试。我认为构建一个这样的基准测试将是一个非常好的社区项目。 最终目标:在一个多样化的真实世界任务集[1]上,根据底层模型和推理工作量进行分组,展示“harness”性能的排行榜(多维度)。任何人都可以贡献结果。 任务标准、测量方法、底层框架等可以由一个小组而非个人决定。 如果兴趣足够,我将创建一个Discord服务器。 披露:我是名为Dirac(https://github.com/dirac-run/dirac)的编码代理的维护者,因此我不会影响最终基准测试的外观,以避免任何利益冲突。我只是想让这件事发生。 [1] 多样化的真实世界任务是指贡献者遇到的足够复杂的任务,最好来自开源仓库。
查看原文
There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one.<p>End goal: a leaderboard of harness performance (multiple axis) on a set of diverse real world tasks[1], grouped by underlying models and reasoning efforts. Anyone can contribute results.<p>The task criteria, measurements, underlying framework et al can be decided by a group rather than a single person.<p>If there is sufficient interest, I will create a discord.<p>Disclosure: I am the maintainer of a coding agent called Dirac (https:&#x2F;&#x2F;github.com&#x2F;dirac-run&#x2F;dirac) so I will not influence what the final benchmark should look like to avoid any conflict of interest. I just want to make this happen.<p>[1] Diverse real world tasks meaning sufficiently complex tasks that the contributors have encountered, preferably from an opensource repo.