Ask HN:更智能的AI模型是如何制造的?

3 分•作者: superasn•14 天前
我很好奇两代人工智能模型之间究竟发生了什么。<p>例如,如何从 Sonnet 发展到 Opus?Opus 是从头开始训练的,还是基于 Sonnet 构建的,抑或是同一个模型但使用了更多的计算资源和训练?<p>像 Astra 这样的模型在某些能力上是如何突然取得巨大飞跃的?是什么阻止了 Anthropic、Mistral 或其他公司做同样的事情?<p>主要区别仅仅是更多的计算资源和资金,还是存在竞争对手可能不知道的训练方法、数据、架构和研究突破?<p>我想不到比这里更好的地方来问这个问题了。我猜这里有人实际上在从事这些模型的开发工作,并且了解幕后发生的事情。
查看原文
I am curious what actually happens between two generations of AI models.<p>For example, how do you go from Sonnet to Opus? Is Opus trained from scratch, built on Sonnet, or mostly the same model with more compute and training?<p>And how do models like Astra suddenly make a big jump in some capabilities? What is stopping Anthropic, Mistral, or others from doing the same thing? Is the main difference just more compute and money, or are there training methods, data, architecture, and research breakthroughs that competitors may not know about?<p>I can&#x27;t think of a better place to ask this. I am guessing there are people here who actually work on these models and know what goes on behind the scenes.