3 分•作者: kenforthewin•3 个月前
返回首页
最新
1 分•作者: pHequals7•3 个月前
1 分•作者: pvdebbe•3 个月前
1 分•作者: mstruebing•3 个月前
1 分•作者: documentorium•3 个月前
1 分•作者: ColinWright•3 个月前
1 分•作者: prokoudine•3 个月前
1 分•作者: tomalpha•3 个月前
1 分•作者: dreampeppers99•3 个月前
1 分•作者: mhb•3 个月前
1 分•作者: Sam6late•3 个月前
4 分•作者: honungsburk•3 个月前
我最近一直在研究爬虫架构。我发现最有用的两个资料是博客文章“2025年,在短短24小时内爬取十亿网页”和Mercator论文(“Mercator:一个可扩展的、可扩展的Web爬虫”)。
这两篇文章,以及我遇到的其他大多数资料,都侧重于爬取广阔的开放网络,而不是针对特定的一组域名。对于产品价格来说,我们关注的是后者。例如,Mercator指出了DNS解析是一个主要的瓶颈,但当你只访问几百个域名时,这实际上并不是一个问题。
另一个差距是,这两者都假定是静态HTML。对于我们的用例,我们需要一个无头浏览器,并且我们还需要处理Cloudflare和类似的防机器人系统。
特别是对于产品价格,许多网站发布价格订阅,这简化了事情,但也有很多网站没有这样做,并且要获得良好的覆盖范围仍然需要抓取。我们目前的系统每天大约处理5亿个页面,我们正在寻求提高其性能。
这里有人在这个领域有经验吗,或者知道关于使用无头浏览器扩展定向(而不是广泛)爬虫的文章/博客文章吗?欢迎提供任何建议。
23 分•作者: ingve•3 个月前
38 分•作者: dberhane•3 个月前
1 分•作者: AndrewVos•3 个月前
1 分•作者: giuliomagnifico•3 个月前
1 分•作者: RivoLink•3 个月前
1 分•作者: sapog_kun•3 个月前
1 分•作者: mohrashard•3 个月前
Built this after watching too many founders get quoted $40k for something
that takes 72 hours to build.<p>You describe your app idea in plain English. Gemini AI breaks it down into
real technical components and generates a realistic agency cost estimate
with timeline and team size. Then you can see the lean MVP alternative.<p>No signup required to see the agency estimate. Curious whether the cost
estimates feel accurate to people who have actually been through this —
and whether the technical breakdowns match what you have seen in the wild.<p>Stack: Next.js 15, Gemini API, Supabase, Resend.
1 分•作者: pseudolus•3 个月前