Ask HN:是否有适用于大型语言模型的优秀安全基准?
2 分•作者: melvinroest•3 个月前
我本人也在寻找这方面的信息,但觉得就此展开实际的讨论会很有益。我对大型语言模型的基准测试方面还比较新。
例如,我看了 eyeballvull [1]。它看起来很有前景,但我没有看到广泛的支持。我尊重作者仍然在坚持开发。我正在寻找一个能够让代理完全扫描代码库的基准测试。
但随后我也在想:也许还有其他我尚未了解的。
而且,冒着可能信息过载的风险,对于那些对具有安全意识的软件工程(无论是否使用代理)感兴趣的人,让我们来聊聊吧!我的邮箱在我的个人资料里。
[1] https://arxiv.org/abs/2407.08708
查看原文
I'm looking for this myself but figured it's good to have an actual discussion about this. I'm pretty new to the benchmarking side of LLMs.<p>For example, I looked at eyeballvull [1]. It seems promising but I don't see wide support for example. I respect the author for still committing. A benchmark where an agent scans a repo in full is what I'm looking for.<p>But then I also wondered: maybe there are others out there that I haven't been aware of yet.<p>And, at the risk of potentially being flooded, for people that are curious about security aware software engineering with agents (or without them for that matter), let's have a chat! My email is in my profile.<p>[1] https://arxiv.org/abs/2407.08708