展示 HN:Echo – 使用开放权重模型,以 1/3 的成本获得寓言级(Fable-level)的成果
35 分•作者: adam_rida•2 个月前
我一直在构建 Echo(<a href="https://echo.tracerml.ai/" rel="nofollow">https://echo.tracerml.ai/</a>),这是一个实验性项目,旨在从一组开源模型中构建一个 AI 系统,而不是选择一个单一模型并将其用于所有任务。
它始于一个简单的实验。我选取了一组模型,包括 GLM-5.2、Kimi K2.7 等,并在相同的评估上运行它们。然后,我测量了如果对于每个问题,你事先知道哪些模型会很有用以及如何组合它们的输出来进行处理,会发生什么情况。
那个假设的系统比池中的任何单个模型都表现得好得多。当然,这并不是一个你可以实际部署的系统,因为它依赖于在看到结果后才知道哪些决策是好的。Echo 是我试图在不事先了解信息的情况下,恢复一部分这种优势的尝试。
对于每个请求,Echo 会决定分配多少计算资源,哪些模型应该参与,以及如何组合它们的工作。有些提示可能只需要相对较少的推理,而有些则受益于多个模型处理问题的不同部分。
在构建它的过程中,令我惊讶的一点是模型之间的互补性。一个整体上明显较弱的模型,在特定问题上或作为组合的一部分仍然可能非常有用。
在我第一次评估混合时,Echo 的表现一直优于其模型池中的最佳单个模型。它还达到了与 Fable 大致相同的聚合结果(我将其用作一个较强的对比系统之一),但推理成本却只有其三分之一左右。
仍然存在一些 Echo 做出错误分配或组合决策的情况。我目前花了很多时间来理解这些失败,以及测试在编码和代理任务上这种方法是否仍然有效,因为在这些任务上衡量每个决策的质量变得更加困难。
我构建了一个聊天界面(echo.tracerml.ai)和一个兼容 OpenAI 的 API(<a href="https://echo.tracerml.ai/docs/api" rel="nofollow">https://echo.tracerml.ai/docs/api</a>),以便系统可以在评估设置之外进行测试。
这是一个关于它如何工作的简短/高层视频:<a href="https://www.youtube.com/watch?v=lJFJSvOdXhg" rel="nofollow">https://www.youtube.com/watch?v=lJFJSvOdXhg</a>
我在这里写下了评估方法、单个模型结果、成本和当前限制:<a href="https://echo.tracerml.ai/eval" rel="nofollow">https://echo.tracerml.ai/eval</a>
我非常希望你能尝试一下!特别是如果你遇到任何奇怪的失败案例或分配看起来不直观的地方。
查看原文
I’ve been building Echo (<a href="https://echo.tracerml.ai/" rel="nofollow">https://echo.tracerml.ai/</a>), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task.<p>It started with a simple experiment. I took a group of models, including GLM-5.2, Kimi K2.7 and others, and ran them on the same evaluations. Then I measured what would happen if, for each problem, you somehow knew in advance which models would be useful and how their outputs should be combined.<p>That hypothetical system performed substantially better than any individual model in the pool. Of course, it is not something you can actually deploy because it relies on knowing which decisions were good after seeing the result. Echo is my attempt to recover some of that advantage without having that information in advance.<p>For each request, Echo decides how much computation to allocate, which models should participate, and how their work should be combined. Some prompts may only need a relatively small amount of inference, while others benefit from multiple models working on different parts of the problem.<p>One thing that surprised me while building it was how complementary the models are. A model that is clearly weaker overall can still be extremely useful on particular problems or as part of a combination.<p>On my first evaluation mix, Echo consistently performed better than the best individual model in its pool. It also reached roughly the same aggregate result as Fable, which I used as one of the stronger comparison systems, at around one third of the inference cost.<p>There are still some cases where Echo makes the wrong allocation or combination decision. I’m currently spending a lot of time understanding those failures, as well as testing whether the same approach holds up on coding and agentic tasks where measuring the quality of each decision becomes much harder.<p>I built a chat interface (echo.tracerml.ai) and an OpenAI-compatible API (<a href="https://echo.tracerml.ai/docs/api" rel="nofollow">https://echo.tracerml.ai/docs/api</a>) so the system can be tested outside the evaluation setup.<p>Here is a short/high level video on how it works: <a href="https://www.youtube.com/watch?v=lJFJSvOdXhg" rel="nofollow">https://www.youtube.com/watch?v=lJFJSvOdXhg</a><p>I wrote up the evaluation methodology, individual model results, costs and current limitations here: <a href="https://echo.tracerml.ai/eval" rel="nofollow">https://echo.tracerml.ai/eval</a><p>I would love for you to try it! Especially if you hit any weird failure cases or places where the allocation looks unintuitive.