HN 提问:您如何保护自托管网站免受 LLM 爬虫的侵害?
1 分•作者: atmosx•大约 1 个月前
您好 HN,
我正在考虑在一台虚拟机上自托管一些流量较低的静态网站。我想了解大家是如何保护个人和专业网站免受滥用性 LLM(及其他)爬虫的侵害的。假设不使用 CDN。
我了解以下方法:
* (1) Anubis AI 防火墙
* (2) LLM 投毒项目,如 iocaine
* (3) 隐藏在 HTML 源代码中的诱饵链接,访问时会屏蔽 IP
* (4) 屏蔽所有主要云服务提供商(GCP、AWS、Azure、Alibaba 等)的 CIDR 范围。
还有其他值得关注的技术吗?
谢谢!
查看原文
Hello HN,<p>I'm considering self-hosting a couple of static low-volume traffic on a VM. I'm curious to know what techniques you using to protect personal and professional websites from abusive LLM (and other) scraping. Assume no CDN is involved.<p>I'm aware of:<p><pre><code> (1) anubis AI firewall
(2) LLM poisoning projects like iocaine
(3) bate links hidden in the HTML source code that block IPs when accessed
(4) ppl blocking CIDRs of all major Cloud Providers (GCP, AWS, Azure, Alibaba, etc).
</code></pre>
Are there any other interesting techniques that can be used?<p>Ty!