DGX Spark 新推理服务器:大型模型 C4:55-90 token/秒,无特定解码
3 分•作者: medicis123•2 个月前
大家好,我们非常激动地与大家分享我们在专为 DGX Spark 集群运行多模型智能体工作流而构建的新推理服务器上的性能数据和基准报告。我们在 2 个 DGX Spark 集群配置上运行了 LlamaBench 测试以及我们自己的模拟流量测试,并获得了一些非常出色的数据。详情如下。完整的详细报告可在我们的 WoolyAI 网站上找到。
WoolyAI 私有 DGX Spark 多智能体推理栈旨在使企业内部团队能够为业务智能体工作流应用程序设置自己的私有、低成本推理栈。它的构建基于一个前瞻性的愿景,即公司将需要自己的私有、低成本推理设置,并且基于单一模型的推理不足以应对复杂的企业工作流智能体应用程序。这些工作流需要不同专业化(因此也不同大小)的多个模型来处理工作流中的不同步骤。为每个模型配备专用的多 GPU 推理栈成本过高。
第一个基准测试:我们在推理服务器上,在 3 个模型上,于 2 个 DGX Spark 集群上运行了 LlamaBenchy 测试,未进行量化和投机解码。使用投机解码将产生更高的数据。
| 模型 | LOAD-PREFILL-SYSTEM-TOK/S | DECODE-SYSTEM-TOK/S | DECODE-TOK/S-PER-REQUEST |
| :----------------------- | :------------------------ | :------------------ | :----------------------- |
| DeepSeek V4 Flash C1 | 1,518.91 | 21.15 | 21.15 |
| DeepSeek V4 Flash C4 | 1,533.15 | 55.99 | 14.00 |
| Gemma 4 26B A4B C1 | 4,579.73 | 30.22 | 30.22 |
| Gemma 4 26B A4B C4 | 4,702.16 | 63.75 | 15.94 |
| Nemotron 3 Nano Omni 30B NVFP4 C1 | 2,607.86 | 39.42 | 39.42 |
| Nemotron 3 Nano Omni 30B NVFP4 C4 | 2,588.87 | 90.83 | 22.71 |
第二个基准测试:一个端点,三个不同的模型控制激活。调度器批量处理每个突发请求,协调两个 rank,并且仅在安全边界处更改驻留模型。
| 模型 | PREFILL-TOK/S | DECODE-TOK/S | Model-Activation-Wait |
| :----------------------- | :------------ | :----------- | :-------------------- |
| DeepSeek V4 Flash | 4,154.34 | 49.30 | 16s |
| Gemma 4 26B A4B | 4,781.44 | 64.67 | 6s |
| Nemotron 3 Nano Omni 30B NVFP4 | 2,395.59 | 93.31 | 2s |
我们认为通过进一步优化,我们可以将这些数据提高 20%。请分享您的反馈。https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/
查看原文
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website.<p>WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.<p>First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.<p>MODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST<p>DeepSeek V4 Flash C1 1,518.91 21.15 21.15<p>DeepSeek V4 Flash C4 1,533.15 55.99 14.00<p>Gemma 4 26B A4B C1 4,579.73 30.22 30.22<p>Gemma 4 26B A4B C4 4,702.16 63.75 15.94<p>Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42<p>Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71<p>Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary<p>MODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait<p>DeepSeek V4 Flash 4,154.34 49.30 16s<p>Gemma 4 26B A4B 18 4,781.44 64.67 6s<p>Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s<p>We think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/