我在消费级硬件上测试了一些本地大语言模型,以下是一些发现。

1 分•作者: felineflock•大约 1 个月前
我使用 LM Studio 在配备 24GB RAM 的 Mac M4 Pro 上对本地 LLM 进行了基准测试。我主要测试了 4 位量化模型,包括 MLX 和 GGUF,模型大小从 4b 到 35b 不等,速度在 3 到 40 tokens/秒之间。 简要结果: - 快速小型模型 -> 提取/分类 - Gemma -> 摘要 - gpt-oss -> 文本咨询 - 大型 Qwen -> 复杂推理/代码解释 没有一个 LLM 能胜任所有任务。 最佳摘要器: 我提供了一段文本,让 LLM 进行摘要,然后对摘要进行了评分。 Gemma 4 e4b 是最佳模型,仅耗时 30 秒。qwen3.6-35b-a3b-ud q2_K_XL 耗时约 2 分钟,并出现了一些遗漏。 最佳“文件问答”(RAG): 我添加了一些文本文件,并就文件中应包含的事实提问。 gpt-oss-20b 在约 4 秒内完美回答了问题。BTL 4 Compact 和 Gemma 4 e4b 也做到了,但耗时约一分钟。 最佳 PII 识别器: 我要求识别文本文件中的人名、组织名以及 PII(电子邮件、电话等)。 Qwen 3 4B MLX(非思考模型)耗时不到 3 秒。其他模型也表现良好:BTL 4 Compact、Gemma 4 e4b、Qwen 3.5 9b 4-bit,但耗时在 11 秒到 30 秒之间。 最佳“代码解释器”: 我提供了一个类文件,并要求 LLM 识别特定的方法和结果值,并判断代码是否可以编译和运行。 Qwen 3.8 27B 在复杂的代码解释方面表现最佳,但耗时 30 分钟(约 5 tok/秒)。BTL 4 Compact 和 Gemma 4 e4b 耗时约一分钟,并出现了一些错误。 我没有测试代码生成或其他任务。 在这些示例中,更多的思考并没有带来更好的答案。有些模型花费了 45 分钟,仍然犯了错误。 Qwen 3.8 27B 可能需要一个推理预算才能在 24GB RAM 上使用(它会不断消耗上下文进行思考,直到我尝试了 1k tokens 的预算)。 有时 Q2 的 Qwen 3.6 变体可以胜过 Q3 变体。 我对其他机器上的可比测试结果很感兴趣,特别是 Mac Studio、AMD Strix Halo、DGX Spark 和消费级 NVIDIA GPU。 你们的模型和工作负载运行得怎么样?
查看原文
I have been benchmarking local LLMs on a Mac M4 Pro 24 GB RAM using LM Studio. I&#x27;ve tested mostly with 4-bit quantization, both MLX and GGUF, from 4b to 35b models, with speeds of 3 to 40 tokens&#x2F;second.<p>Results briefly:<p>- fast small model -&gt; extraction&#x2F;classification<p>- Gemma -&gt; summarization<p>- gpt-oss -&gt; transcript consultation<p>- large Qwen -&gt; difficult reasoning&#x2F;code interpretation<p>There wasn&#x27;t a single best LLM for all tasks.<p>Best summarizer: I took a transcript and asked an LLM to summarize, then graded the summarization. Gemma 4 e4b was the best and took only 30s. qwen3.6-35b-a3b-ud q2_K_XL took around 2 minutes and incurred in a few omissions.<p>Best &quot;answerer from files&quot; (RAG): I added some transcript files and asked questions about facts that should be grounded on the files. gpt-oss-20b answered questions perfectly in about 4s. BTL 4 Compact and Gemma 4 e4b did too but took around a minute.<p>Best PII identifier: I asked to identify names of people and organizations plus PII (emails, phones, etc) in a text file. Qwen 3 4B MLX (not a thinking model) took less than 3s. Others were good too: BTL 4 Compact, Gemma 4 e4b, Qwen 3.5 9b 4-bit - but they took between 11s and 30s.<p>Best &quot;code interpreter&quot;: I gave a class file and asked an LLM to identify certain methods and result values and tell whether the code would compile and run. Qwen 3.8 27B was by far the best at difficult code interpretation, but at the cost of 30 minutes (~5 tok&#x2F;s). BTL 4 Compact and Gemma 4 e4b took around a minute and responded with a couple of mistakes.<p>I did not test code generation or any other types.<p>In these examples more thinking didn&#x27;t result in better answers. Some models spent 45 minutes and still got something wrong.<p>Qwen 3.8 27B may need a reasoning budget to make it usable in 24Gb RAM (it was exhausting the context with thinking over and over until I experimented with a budget of 1k tokens).<p>Sometimes a Q2 Qwen 3.6 variant can beat a Q3 variant.<p>I would be interested in comparable test results from other machines, especially Mac Studio, AMD Strix Halo, DGX Spark, and consumer NVIDIA GPUs.<p>What models and workloads are working well for yall?