Aquila 智能语音助手家庭助理测试套件

1作者: aquila4162 天前
构建一个全面的开源语音助手测试套件。您可以在以下网址查看,包括当前的排行榜:https://git.cicero.sh/aquila/ha-voice-test-suite/ 测试是可复现的,并且有清晰的说明,教您如何在自己的机器上运行它们。 我只有一块 4GB 显存的 GPU,因此只能测试像 Qwen3 4B Instruct 这样的小型 LLM。我现在正在运行 Gemma 4,但速度非常慢,可能还需要 24 到 48 小时才能完成。 我尝试过 Claude Sonnet 5 和 Grok 4.5 等云端模型,但每次都被速率限制,我感到厌烦,所以放弃了。也许有人比我更频繁地使用 AI,并且拥有合适的速率限制,不介意在一些前沿模型上运行测试套件,因为我猜想大家会乐于看到结果。我已有信心,它们在 60-90 分钟的测试时长下能达到 90% 以上的得分,所以对此并不过于担心。 如果有人拥有更大的 GPU,并且不介意在 20B+ 的大型本地 LLM 上进行测试,那就太棒了。在任何模型上运行测试都非常简单,说明在 README 文件中。 祝您使用愉快!
查看原文
Put together an extensive open source test suite for voice assistants. You can view it including current leaderboard at: https:&#x2F;&#x2F;git.cicero.sh&#x2F;aquila&#x2F;ha-voice-test-suite&#x2F;<p>Tests are reproduceable, with clear instructions on how to run them on your machine there.<p>I only have a GPU with 4GB vRAM, hence only capable of testing the small LLMs like Qwen3 4B Instruct. Have Gemma 4 running right now, but it&#x27;s insanely slow and probably another 24 - 48 hours before it finishes.<p>Tried cloud models like Claude Sonnet 5 and Grok 4.5, but was rate limited each time, and got fed up so dropped it. Maybe someone out there uses AI more than me and has proper rate limits and wouldn&#x27;t mind running the test suite on some frontier models as I&#x27;m assuming folks would appreciate seeing the results. I&#x27;m already confident they&#x27;ll get low 90s with 60 - 90 minute duration, so not overly worried about it.<p>If anyone has larger GPU and wouldn&#x27;t mind giving it a spin on larger 20B+ local LLMs that would be awesome. Running the tests against any model is quite straight forward, instructions in the readme.<p>Enjoy!