Launch HN:Speko (YC S26) – 语音 AI 的 OpenRouter

10 分•作者: abdik•大约 1 个月前
大家好!我是 Bek,Speko 的创始人。Speko 是一个平台,可以根据您的约束条件,在我们所有公开基准测试的选项中,找到语音转文本 (STT)、大型语言模型 (LLM) 和文本转语音 (TTS) 模型的最佳组合,并解释原因。 演示:[https://youtu.be/no2LY2gRh-c](https://youtu.be/no2LY2gRh-c) 典型的生产语音代理由三个模型组成:STT、LLM 和 TTS。 每个层都有十几个信誉良好的供应商,而且每个月都有新模型上市。几乎所有人都会评估一次,选择自己喜欢的堆栈,然后不再重新检查,因为从一个供应商切换到另一个供应商需要进行另一次集成,并且会就数据进行争论。 结果是,您使用的语音代理运行的是上个季度的模型,而实际上有更好、更便宜的选项可用。 在创立 Speko 之前,我曾担任联合创始人兼首席技术官,花了四年时间为亚洲的企业构建了十多种语言的语音代理。每次有新的语音模型出现,我们都会重复同样的仪式:雇佣母语为母语的评分员,将其与我们现有的堆栈进行基准测试,如果有所改进就更新生产环境。Speko 将这个过程变成了一个 API。一个每天处理数千次通话的团队告诉我们:“我们可以直接进入这个仪表板,切换模型,它就会为我们完成所有工作。” 工作原理:您发送一个包含优化标准(准确性、延迟、成本或平衡)、语言和地区的请求。路由器会筛选出我们针对给定约束条件组合进行过测量的模型,对其进行基准测试,选择最佳模型,并返回一个响应,其头部包含提供商、模型名称和分数。网关会预取签名的会话计划,因此新会话会直接从内存中拨打提供商;在呼叫者等待时,不会有控制平面往返。 故障转移仅发生在连接建立阶段:如果提供商拒绝连接尝试,我们会开始连接到备选方案。 一些客户案例:一位创始人找到我们时完全不知道该如何选择:他告诉我们他的用例,现在所有流量都通过该平台进行路由。一个物业管理 AI 使用 Python 运行 LiveKit,自推出以来一直没有更新 STT 或 TTS:他们不知道他们的 STT 在通话中存在很高的错误率,存在更好的选项,而更换总是看起来像一个研发项目。一个团队不知道该为西班牙语选择哪些模型。一个医疗团队不知道哪个 STT 最能处理医学词汇。在每种情况下,我们都帮助他们从基准测试中找到了合适的堆栈,现在他们通过我们进行路由。 测量部分是公开的:我们在不同日期对同一地区的所有模型使用相同的输入,并发布排行榜,包括我们选择的模型表现不如替代方案的情况。一个发布演示回答了哪个 30 秒的片段听起来更好;生产环境则询问哪个模型能撑过第八分钟,因此我们测试了即兴语音、金钱和日期、十分钟的录音,排名也会随之改变。我们训练了一个自动评分器来评估 TTS 的自然度,基于我们盲测的正面交锋听力投票;对于它从未见过投票的提供商,它选择的获胜者与我们的评分员选择的获胜者一致的频率,与评分员之间的一致性大致相同。 我们自己不训练或销售模型,这正是我们保持排名公正的原因。 我们还开源了网关,供希望避免音频路径上额外网络跳数且不想与我们的云共享密钥的团队使用([https://github.com/SpekoAI/gateway](https://github.com/SpekoAI/gateway),MIT):一个 Go 二进制文件,作为 sidecar 在您的代理容器中运行,通过 Unix 套接字使用本地协议,固定提供商主机并附加您的密钥。在 BYOK 模式下,它根本不会与我们通信。 请注意,匿名、无内容遥测是默认启用的,一个环境变量即可禁用它。 成本:网关和 BYOK 设置将永久免费,我们收取托管路由器和托管密钥的费用,并提供统一账单。自六月下旬开始批量处理以来,外部使用量平均每周增长约 25%,主要集中在发布周。 我很想听听社区的反馈:您现在是如何选择语音模型的,以及是什么让您信任第三方基准测试?
查看原文
Hi HN! I&#x27;m Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constraints, among all our public benchmarked options, and tells you why.<p>Demo: <a href="https:&#x2F;&#x2F;youtu.be&#x2F;no2LY2gRh-c" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;no2LY2gRh-c</a><p>Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS.<p>Each of those layers offers a dozen credible vendors, and each month there are new models on the market. Almost everyone evaluates once, picks a stack of their choice, and never rechecks because switching from a vendor to another involves yet another integration and arguments about the numbers.<p>The result is that you use voice agents running last quarter&#x27;s models while better and cheaper options are available.<p>Before founding Speko, I spent four years as cofounder and CTO building voice agents for enterprises across Asia in 10+ languages. Each time a new speech model would arrive, we repeated the same ritual: hire native-speaking raters, benchmark it against our existing stack, and update production if it improved. Speko turns this process into an API. A team running thousands of calls a day told us: &quot;we can literally go to this dashboard, switch the model, and it will do it for us.&quot;<p>How it works: you send a request with your optimization criteria (accuracy, latency, cost or balanced), language and region. The router filters to models which we measured for the given combination of constraints, benchmarks them, selects the winner, and returns a response with headers containing provider, model names, and the scores. The gateway prefetches signed session plans, so a new session dials the provider straight from memory; no control-plane round trip while a caller waits.<p>Failover happens only during connection setup stage: if the provider refuses the connection attempt, we start connecting to the runners-up.<p>Some of the customer stories: one founder came to us not knowing what to pick at all: he gave us his use case and now routes everything through the platform. A property management AI runs LiveKit in Python and had not updated STT or TTS since launch: they did not know their STT had high error rates on their calls, better options existed, and swapping always looked like an R&amp;D project. One team did not know which models to pick for Spanish. A medical team did not know which STT handles medical vocabulary best. In every case we helped find the right stack from the benchmarks, and now they route through us.<p>The measuring part is public: we pass the same inputs to every model in one region in different dated runs and we publish the boards, including those where our selections perform worse than alternatives. A launch demo answers which 30-second clip sounds better; production asks which model survives minute eight, so we test spontaneous speech, money and dates, ten-minute takes, and the rankings change. We trained an automatic scorer for TTS naturalness on our blind head-to-head listening votes; on providers it has never seen a vote for, it picks the same winner our raters do about as often as raters agree with each other.<p>We don&#x27;t train or sell models ourselves, that&#x27;s precisely how we keep our rankings impartial.<p>We also open sourced the gateway for teams who want to avoid an extra network hop on the audio path and don&#x27;t want to share keys with our cloud (<a href="https:&#x2F;&#x2F;github.com&#x2F;SpekoAI&#x2F;gateway" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;SpekoAI&#x2F;gateway</a>, MIT): one Go binary, which is running as a sidecar in your agent&#x27;s container, speaks one local protocol over Unix socket, pins provider hosts and attaches your keys. In BYOK mode it doesn&#x27;t communicate with us at all.<p>Notice that the anonymous, content-free telemetry is enabled by default, and one env var disables it.<p>Cost: the gateway and BYOK setup will be free forever, we charge for the hosted router and managed keys with consolidated billing. Since we started the batch in late June, external usage has grown about 25 percent per week on average, front-loaded toward the launch weeks.<p>I would love feedback from the community: how do you pick speech models now, and what makes you trust the third-party benchmark?