用于阐释性人工智能的创意

1 分•作者: E-Reverance•大约 1 个月前
我听说了一些关于前沿模型在“解释数学”方面仍然表现不佳的抱怨,我正在考虑一个可能有帮助的强化学习环境,那就是: * 选取非常难但有可验证答案的数学问题。 * 让前沿模型向一个非常小的模型(例如参数量为 0.5-1B 且在此类问题上得分明显偏低)解释如何解决问题,但不要直接给出答案。前沿模型会因为其提示/解释帮助了小模型解决问题而获得奖励。 显然,需要一定程度的人工监督来避免前沿模型提供过多的信息。
查看原文
I&#x27;ve heard some complaints about the frontier models still be bad at <i>explaining math</i> and was thinking of an RL environment that would help might be to:<p>-Take very hard math problem with a verifiable answer<p>-Have frontier model explain to a tiny model like (0.5-1B params and provably bad score on the problem) how to solve but not the solution, and reward the frontier model for prompts&#x2F;explanations that helped the tiny model solve the problem<p>Obviously some amount of human supervision is needed to weed out it giving too much information