展示 HN:我将我的代理 Token 数量减少了 50%
2 分•作者: iohotspot•大约 1 小时前
这一切都始于观察 Claude 代码中内置的工具调用。我们注意到,花费了大量时间使用正则表达式搜索代码行。当你在一个未知的环境中工作时,这种方法是有意义的,但我们都在某种代码库结构和跟踪系统中工作。这让我们深入研究了代码库索引,并促使我们构建了 Indexio。
Indexio 的构建旨在消除代理在重新发现代码库时浪费的时间和 token。每次正则表达式调用都会返回大量数据到上下文窗口,随着时间的推移,这会导致不真正重要的项目膨胀。由于代码库工作主要包含静态文件,因此索引它们并快速获取数据以供发现是我们试图解决的问题。
Indexio 将代码库索引到其内部数据库,然后在被请求时通过 MCP 返回所需内容。我们在尝试强制会话使用工具调用进行查找时遇到了一些有趣的问题,因此存在一些激进的工具调用逻辑,试图覆盖内置的查找工具。如果索引中没有某个内容,它将回退到正则表达式查找。
整个系统是用 Rust 编写的单个二进制文件,以实现高性能。在我们实际测试中,与内置工具相比,token 使用量减少了约 50%-60%,查找时间接近即时,而每次正则表达式查找需要数秒钟。Indexio 在查找符号方面确实不如正则表达式,我们发现大多数代理的工作都是查看单词字符串而不是符号,但如果你有一个在此特定场景中工作量很大的工作流程,请注意它可能不适合。
我们开源了所有内容,并非常希望听到反馈。非常感谢 CocoIndex,他们给了我们这个想法的灵感。
https://github.com/IOServicesLabs/indexio
查看原文
This started by watching the built-in tool calls in Claude code. We noticed that a LOT of the time was being spent searching for lines of code with regex. This approach makes sense when you are working in an unknown environment, but we all work inside of some codebase structure and tracking. This took us down the rabbit hole of codebase indexing and led to us building Indexio.
Indexio was built to eliminate the time and tokens agents waste rediscovering a codebase. Regex calls return lots of data each time to the context window which causes bloat over time with items that don't really matter. Since codebase work container mostly static files indexing them and quickly getting to data for discovery is what we were trying to solve.<p>Indexio index's the codebase to its internal DB and then returns what is needed via MCP when asked. We ran into some fun issues when trying to force the session to use a tool call for lookups so there is some aggressive tool call logic which attempts to override built-in lookup tools. If something isn't in the index it will just fallback to regex lookups.<p>The whole system is a single binary written in Rust for performance. In our real world testing we are seeing about 50%-60% reduction in token usage compared to the built-in tools, and near instant lookup times compared to multiple seconds for each regex lookup. One place that indexio does come up shorter than then regex is when looking up symbols. We found that most of the agent's work was done looking at strings of words rather than symbols, but if you have a workflow that is heavy in this specific situation just know it may not fit.<p>We opened sourced everything and would love to hear feedback. Huge Shout out to CocoIndex who gave us the inspiration for this idea.<p><a href="https://github.com/IOServicesLabs/indexio" rel="nofollow">https://github.com/IOServicesLabs/indexio</a>