HN 提问:我应该结合全局知识、互联网搜索和用户 RAG 吗?

1 分•作者: abdrhxyiii•2 个月前
我正在斯里兰卡构建一个处理文档和其他敏感数据的 SaaS 平台。 每个用户都可以上传自己的文档和信息,平台使用 RAG(检索增强生成)根据该用户的数据来回答问题。这部分我理解。 我主要担心的是,当用户上传的信息不足时会发生什么。我仍然希望 LLM(大型语言模型)能够使用来自互联网(或经过策划的知识库)的可靠信息提供准确的答案,并附带适当的引用。 以下是我正在考虑的两种架构: 选项 1: 基础 LLM(通过 Azure AI Foundry 或 Amazon Bedrock 使用 OpenAI/Anthropic) ↓ 平台 RAG(由我们管理的全局知识库) ↓ 用户特定 RAG 在这种方法中,我们维护一个由我们(平台管理员)策划和更新的全局知识库。每个用户都可以访问这个共享知识,同时通过他们个人的 RAG 搜索他们自己上传的文档。 选项 2: 开源 LLM ↓ 在斯里兰卡/领域特定数据上进行微调 ↓ 用户特定 RAG 在这里,我们使用斯里兰卡或领域特定的数据对开源模型进行微调,并且每个用户仍然拥有自己的 RAG 来处理其私有文档。 我的顾虑是: 微调实际上是正确的解决方案吗,还是不必要的? 全局/共享 RAG 是否比微调更好? 如果您想实现以下目标,您将如何设计此架构: 来自领域知识的准确答案 用户私有文档搜索 引用/来源 为数千用户提供良好的可扩展性 我倾向于选项 1,因为微调似乎成本高昂、耗时,而且我对此还没有经验。但是,我不确定我的想法是否正确。 我非常希望听听其他人会如何处理这个问题。
查看原文
I&#x27;m building a SaaS platform in Sri Lanka that handles documents and other sensitive data.<p>Each user can upload their own documents and information, and the platform uses RAG to answer questions based on that user&#x27;s data. That part makes sense to me.<p>My main concern is what happens when the user hasn&#x27;t uploaded enough information. I still want the LLM to provide accurate answers using reliable information from the internet (or from a curated knowledge base), with proper citations.<p>These are the two architectures I&#x27;m considering:<p>Option 1:<p>Base LLM (OpenAI&#x2F;Anthropic via Azure AI Foundry or Amazon Bedrock) ↓ Platform RAG (global knowledge base managed by us) ↓ User-specific RAG In this approach, we maintain a global knowledge base that we (the platform admins) curate and update. Every user can access this shared knowledge, while their own uploaded documents are searched through their personal RAG.<p>Option 2:<p>Open-source LLM ↓ Fine-tuned on Sri Lankan&#x2F;domain-specific data ↓ User-specific RAG Here, we fine-tune an open-source model using Sri Lankan or domain-specific data, and each user still has their own RAG for their private documents.<p>My concerns are:<p>Is fine-tuning actually the right solution here, or is it unnecessary?<p>Is a global&#x2F;shared RAG a better approach than fine-tuning?<p>How would you design this architecture if you wanted:<p>Accurate answers from domain knowledge<p>User-private document search<p>Citations&#x2F;sources<p>Good scalability for thousands of users<p>I&#x27;m leaning toward Option 1 because fine-tuning seems expensive, time-consuming, and I have no experience with it yet. However, I&#x27;m not sure if I&#x27;m thinking about this correctly.<p>I&#x27;d really appreciate hearing how others would approach this problem