提问 HN:切换到自托管推理的优缺点是什么?
1 分•作者: codenski•4 个月前
在围绕数据隐私进行了一些合规性讨论后,管理层正推动我们在内部运行开源模型。在做出决定之前,我们很希望听取已经完成此项过渡的人的经验。<p>我们特别感兴趣的是:<p>与按照请求量支付 API 访问费用相比,实际成本是否更低?
在管理性能方面,特别是延迟、吞吐量和硬件利用率方面,是否存在任何问题?
您如何处理跨团队/工作负载的成本可见性和归属问题?<p>此外,我们对其他方面也超级感兴趣,例如哪些有效,哪些无效,以及在切换之前您希望了解什么?<p>提前感谢!
附言:我们并非寻求绝对真理,只是希望为可能的过渡做好准备。
查看原文
Management is pushing us toward running open-weight models in-house after some compliance conversations around data privacy. Before we commit, we'd love to hear from people who've made this transition.<p>Specifically curious about:<p>Did it actually end up cheaper than paying for API access at your request volume?
Were there any issues related to managing performance, more specifically latency, throughput, hardware utilization?
How do you handle cost visibility and attribution across teams/workloads?<p>Also, super curious about other aspects, what worked, what didn't, and what do you wish you'd known before switching?<p>Thanks in advance!
PS: We are not seeking for an absolute truth, just want to be prepared if that transition will take place.