HN 提问:我应该结合全局知识、互联网搜索和用户 RAG 吗?
1 分•作者: abdrhxyiii•2 个月前
我正在斯里兰卡构建一个处理文档和其他敏感数据的 SaaS 平台。
每个用户都可以上传自己的文档和信息,平台使用 RAG(检索增强生成)根据该用户的数据来回答问题。这部分我理解。
我主要担心的是,当用户上传的信息不足时会发生什么。我仍然希望 LLM(大型语言模型)能够使用来自互联网(或经过策划的知识库)的可靠信息提供准确的答案,并附带适当的引用。
以下是我正在考虑的两种架构:
选项 1:
基础 LLM(通过 Azure AI Foundry 或 Amazon Bedrock 使用 OpenAI/Anthropic)
↓
平台 RAG(由我们管理的全局知识库)
↓
用户特定 RAG
在这种方法中,我们维护一个由我们(平台管理员)策划和更新的全局知识库。每个用户都可以访问这个共享知识,同时通过他们个人的 RAG 搜索他们自己上传的文档。
选项 2:
开源 LLM
↓
在斯里兰卡/领域特定数据上进行微调
↓
用户特定 RAG
在这里,我们使用斯里兰卡或领域特定的数据对开源模型进行微调,并且每个用户仍然拥有自己的 RAG 来处理其私有文档。
我的顾虑是:
微调实际上是正确的解决方案吗,还是不必要的?
全局/共享 RAG 是否比微调更好?
如果您想实现以下目标,您将如何设计此架构:
来自领域知识的准确答案
用户私有文档搜索
引用/来源
为数千用户提供良好的可扩展性
我倾向于选项 1,因为微调似乎成本高昂、耗时,而且我对此还没有经验。但是,我不确定我的想法是否正确。
我非常希望听听其他人会如何处理这个问题。
查看原文
I'm building a SaaS platform in Sri Lanka that handles documents and other sensitive data.<p>Each user can upload their own documents and information, and the platform uses RAG to answer questions based on that user's data. That part makes sense to me.<p>My main concern is what happens when the user hasn't uploaded enough information. I still want the LLM to provide accurate answers using reliable information from the internet (or from a curated knowledge base), with proper citations.<p>These are the two architectures I'm considering:<p>Option 1:<p>Base LLM (OpenAI/Anthropic via Azure AI Foundry or Amazon Bedrock)
↓
Platform RAG (global knowledge base managed by us)
↓
User-specific RAG
In this approach, we maintain a global knowledge base that we (the platform admins) curate and update. Every user can access this shared knowledge, while their own uploaded documents are searched through their personal RAG.<p>Option 2:<p>Open-source LLM
↓
Fine-tuned on Sri Lankan/domain-specific data
↓
User-specific RAG
Here, we fine-tune an open-source model using Sri Lankan or domain-specific data, and each user still has their own RAG for their private documents.<p>My concerns are:<p>Is fine-tuning actually the right solution here, or is it unnecessary?<p>Is a global/shared RAG a better approach than fine-tuning?<p>How would you design this architecture if you wanted:<p>Accurate answers from domain knowledge<p>User-private document search<p>Citations/sources<p>Good scalability for thousands of users<p>I'm leaning toward Option 1 because fine-tuning seems expensive, time-consuming, and I have no experience with it yet. However, I'm not sure if I'm thinking about this correctly.<p>I'd really appreciate hearing how others would approach this problem