问 HN:潜在空间不仅仅是压缩吗?我们可以探究它的内部规则吗?

1作者: WLHsu3 个月前
我很好奇,人们是否认为Anthropic的注入式思维检测工作与LeWM等潜在状态世界模型之间存在更深层的联系。 一方面,模型似乎能够报告其自身内部扰动的一部分。另一方面,训练被明确地推向一个结构更清晰的潜在预测空间。 这些研究是否表明,潜在空间可能部分是可探测和可解释的,可以被视为一个内部规则空间,而不仅仅是一个压缩的向量空间? 有没有人尝试过结合潜在状态探测、正则化世界模型和内省式检测?
查看原文
I’m curious whether people see a deeper connection between Anthropic’s injected-thought detection work and latent-state world models like LeWM.<p>In one case, the model seems able to report parts of its own internal perturbation. In the other, training is explicitly pushed into a more structured latent prediction space.<p>Do these lines of work suggest that latent space may be partially probeable and interpretable as an internal rule space, rather than just a compressed vector space? Has anyone experimented with combining latent-state probing, regularized world models, and introspection-style detection?