HN 提问:最近一次只有前沿模型才能完成的任务是什么?
14 分•作者: thedebuglife•3 个月前
我看到一种反复出现的说法,即落后前沿模型6个月的开放(权重)模型足以应对大多数“工作”。如果您在过去一个月内遇到一个具体的任务,其中GLM/DeepSeek/Kimi/Qwen失败而Opus/Fable/GPT成功(反之亦然!),请分享。
我将提供一个模板,因为我也经常看到发帖者抱怨缺乏上下文:
- 任务(简短且具体)
- 尝试的廉价模型及其失败原因
- 前沿模型及其是否成功
- 事后看来,前沿-1是否就足够了?
查看原文
ive been seeing a recurring claim that open (weight) models 6 months behind the frontier are good enough for the majority of ‘work’. if you've had a concrete task in the last month where GLM/DeepSeek/Kimi/Qwen failed and Opus/Fable/GPT succeeded (or vice versa!), please share<p>Ill provide a template as ive also frequently seen posters complain about a lack of context:
- task (short and specific)
- cheap model tried & how it failed
- frontier model & did it actually succeed
- would frontier-1 have been fine in hindsight?