Ask HN:您是如何控制 Token 成本的?

1 分•作者: jainojas•3 个月前
我从 2024 年初开始使用大型语言模型(LLM)和编码助手。总的来说,编码助手和大型语言模型面临的一个大问题是上下文压缩。 举例来说,当我分析自己在使用 Claude Code、Codex 和 Sakana 时的会话时,我发现我的大部分助手花费了超过 90% 的时间在重复阅读上下文。进一步研究它们阅读的 Markdown 文件后,我粗略估计其中至少有 20% 的内容对当前任务毫无用处。 深入研究这个问题后,我意识到这是一个前沿研究领域。有人甚至提出了解决方案,例如让大型语言模型使用一种人类无法理解的、抽象的压缩语言进行推理,这种语言比人类语言更节省 token,然后再使用一个解码器模型来保证人类的可读性和可访问性。 我想了解一下,目前还有哪些其他方法被应用?您在使用这些助手时有什么经验?您是否也像我一样担心这种“token 损耗”?
查看原文
I have been using LLMs &amp; Coding Agent since early 2024. A large problem with Coding Agents &amp; LLMs in general is context compression.<p>To give you some numbers, when I analysed my own sessions across Claude Code, Codex &amp; Sakana, I found that most of my agents spent &gt;90% of time re reading context and upon further investigation into the markdowns it was reading, I have a hand-wavy estimate of at least ~20% of this being useless to the task at hand.<p>When digging a bit more into this problem, I realised that this is an active area of frontier research, wherein some have even proposed solutions like having the LLM reason in an abstract compressed language illegible to humans which is more token efficient than human languages &amp; then using a decoder model on top of this for human readability &amp; access.<p>Curious to know, what other approaches are being used out there ? What is your experience of working with these agents &amp; are you concerned about this &quot;token-rot&quot; as I call it or not ?