技能 x2,108 次运行:优化功效和代币效率

1 分•作者: darvh•大约 1 个月前
马尾辫 vs Signal 基准测试虽然运行起来了,但过程有些混乱。模型表现出了对答案的记忆,这在如今也是意料之中的。 比起技术本身,更令人开心的是成功运行了基准测试 :)。 如果关于基准测试的进行方式或构建方式的介绍有任何问题,请告诉我。 博客文章:darvh.com/posts/when-coding-agents-raced-through-108-bugs/ 基准测试:github.com/darvh/bench Signal:github.com/darvh/signal
查看原文
Ponytail vs Signal<p>Bench has the runs albeit not in an organised manner. Models showed memorization of the answers, which is kinda expected these days.<p>More than the skills, it was fun getting the benchmark running :).<p>Let me know if the preamble on how the bench was conducted or constructed is flawed.<p>Blog Post: darvh.com&#x2F;posts&#x2F;when-coding-agents-raced-through-108-bugs&#x2F;<p>Bench: github.com&#x2F;darvh&#x2F;bench<p>Signal: github.com&#x2F;darvh&#x2F;signal