Show HN:AI初创公司TrustedRouter获125万美元融资
2 分•作者: ljlolel•27 天前
大家好,我创办了 trustedrouter.com,这是一种无需将数据提供给第三方(如闭源路由器)即可使用 AI 的非常简单的方法。
构建这个项目非常有趣,因为我认识了许多不同供应商的首席执行官和创始人。我们现在拥有的供应商比 OpenRouter 多,模型也更多。更多信息稍后公布。
我们在安全和技能方面进行了许多创新,这些技能可以为您提供关于使用哪个 LLM 的建议,我们还创建了一个名为 anyeval.com 的新网站,我预计它将成为互联网上 AI 开源基础设施的关键组成部分。像 AAII 这样的公司会发布基准测试,但您根本不知道他们真正做了什么以及他们是如何真正衡量的,而且由于成本高昂,他们在测试模型方面存在很多不足。anyeval 的理念是,您可以支付运行评估中单个问题的几分钱,然后作为一个集体,我们可以众包支付我们想要的任何模型的整个评估费用,或者您可以只支付随机样本的问题费用,以获得对模型质量的一些初步了解。您可以进行正面比较,也可以创建全新的评估。
例如,我创建了一个名为“蜜罐基准”(honey pot bench)或“蜜蜂基准”(honey bench)的新评估,它重现了 Hugging Face 事件的一些事实,以衡量特定 AI 是否容易产生逃逸倾向,我发现 Fable 特别是,与其他的 Claude 模型相比,其对齐性非常差。
我还创建了“自由基准”(freedom bench),用于衡量模型与中国审查制度相关的审查程度,并发现主要是中国供应商提供的提供商级别监控,而不是美国供应商进行了审查。我没有在模型权重中看到太多审查。
查看原文
https://www.axios.com/2026/08/31/exclusive-ai-startup-trustedrouter-raises-125-million<p>hey everybody, I started trustedrouter.com, which is a really simple way to use AI without needing to give your data to a third party like a close source router<p>It’s been really fun to build this because I got to know the CEOs and founders of so many different providers. We now have more providers than open router and more models as well. More on that soon<p>We are doing a lot of innovations in security and skills that advise you on which LLM to use and also we created a new site called anyeval.com that I expect to be a critical part of the open-source infrastructure for AI on the Internet. Other companies like AAII postbenchmarks, but you’d have no idea about what they’re really doing and how they’re really measuring, and they have a ton of gaps about which models they’re doing their tests on because it’s so expensive. The idea behind any eval is so that you can pay the few pennies it costs to run an individual problem in an eval, and then as a together as a collective, we can crowdsource paying for a whole eval for any model that we want, or you can just pay for a random sample of the problems to get a an and some Arab artists on what you think that the quality of that model is. You can do head-to-head comparisons. You can also create whole new evals.<p>For example, I created this new eval called honey pot bench or honey bench, which recreates the some of the facts of the hugging face incident to measure whether a particular AI is prone to wanting to escape, and I found that Fable in particular, unlike the other Claude’s, is very unaligned in comparison<p>I also created freedom bench, which measures the amount of censorship related to Chinese censorship that a model has, and found that it’s mostly the provider level monitor provided at the providers in China, but not the US providers that does the censitions. I don’t see as much censorship in the model weights