数据是最后的护城河

1作者: yckanishk5 个月前
事情是这样的:如果因为大语言模型(LLM)之间相互竞争,都想击败对方,导致智能被商品化,价格不断下降,感觉这里正在发生一场价格战。如果这种智能变得如此便宜,就像 Sam Altman 说的那样,便宜到可以像计量用电一样计量,然后如果 LLM 代理(比如 openclaw 或 hermes)的“躯体”,也就是框架,开源了,并且比任何初创公司提供的都好,那么数据就成为了最后的护城河。归根结底,每个超级智能或通用人工智能(AGI)实际上都只擅长处理某些子集任务。如果真是这样,那么影响它们训练的数据,如果预训练持续进行,没有零训练的 LLM 出现,那么影响这些 LLM 的数据将成为决定该 LLM 擅长哪些子集任务的主要因素。因此,数据就成为了最后的护城河。对此有什么看法?(这是我在 Hacker News 上的第一篇文章 :))
查看原文
Here is the thing: if intelligence gets commoditized because LLMs are competing with each other, each trying to beat each other, and the prices are going down and down, it feels like there is a pricing war going on here. If this intelligence is now available so cheaply, like how Sam Altman said that it will be so cheap that it can be measured like electricity is metered, and then if the body, the harness an that makes an LLM agent like openclaw or hermes, becomes open source and it becomes better than anything that a startup can provide, then data becomes the final moat. At the end of the day every super intelligence or AGI will actually be good at only a subset of things. If that is true then the data that is affecting their training, if the pre-training thing continues and no zero-training LLMs are produced, then what happens is that the data that is affecting these LLMs will become the primary factor in determining which subset of tasks that LLM is good at. Therefore data becomes the final moat. any thoughts on this? (this is my first post on hackernews :))