Ask HN:HotPin – 在 24GB RAM 上实现无损 120B MoE 推理(CPU,50 行代码)

2 分•作者: LozzKappa•2 个月前
我是一名机电一体化设计师,拥有控制系统、机器人、PCB设计和嵌入式硬件的背景。我设计物理系统:电机、传感器、微控制器和实时控制回路。 我将这种设计思维应用于LLM内存管理,并且奏效了。 HotPin是一组针对llama.cpp的补丁,它可以在远少于模型磁盘占用的RAM下运行30B-120B的混合专家(MoE)模型,并实现位精确(无损)输出。 在AMD Ryzen AI 9 HX 370 (Zen5, AVX512)、23.6GB LPDDR5X、NVMe >1GB/s、仅CPU环境下进行测试。 结果: | 模型 | 磁盘 | 最低RAM | 节省 | tok/s | |---|---|---|---|---| | gpt-oss:120b | 58.5GB | 19.1GB | -67% | 3.84 | | qwen3:30b-a3b | 18.0GB | 10.4GB | -42% | 19.7 | | gemma4:26b-a4b | 16.2GB | 10.6GB | -35% | 11.5 | | GLM-4.7-Flash | 19.0GB | 13.3GB | -30% | 12.4 | 输出与全RAM运行的SHA-256哈希值位精确相同。已验证。 工作原理(llama.cpp中约50行C++代码): 1. 分析MoE专家路由频率。 2. 从磁盘mmap整个模型。 3. 将最热的专家mlock到物理RAM中。 4. 在需要之前,使用posix_fadvise / prefetch将冷专家从NVMe预加载。 边界条件:如果磁盘 > RAM,固定(pinning)可带来+45%的速度提升(gpt-oss:2.64 → 3.84 tok/s)。如果模型适合RAM,固定不会增加任何开销。 测试平台:Linux原生、WSL、Windows原生(VirtualLock)。 代码库:https://github.com/LozzKappa/hotpin-llm 论文(PDF + LaTeX)在代码库中——arXiv提交正在等待cs.LG领域的背书。 我正在寻找: 1. 一位cs.LG领域的arXiv背书人——如果您是研究人员并且对这项工作感兴趣,请联系我。 2. 任何在自己硬件上测试并提供反馈的人。 该技术简单、无损,并且即刻可用。请测试它并告诉我您的基准测试结果。 感谢阅读。
查看原文
I&#x27;m a mechatronics designer with a background in control systems, robotics, PCB design, and embedded hardware. I design physical systems: motors, sensors, microcontrollers, and real-time control loops.<p>I applied this design thinking to LLM memory management – and it worked.<p>HotPin is a set of patches for llama.cpp that runs 30B–120B Mixture of Experts (MoE) models on far less RAM than their disk footprint, with bit-identical (lossless) output.<p>Tested on an AMD Ryzen AI 9 HX 370 (Zen5, AVX512), 23.6GB LPDDR5X, NVMe &gt;1GB&#x2F;s, CPU-only.<p>Results: | Model | Disk | Min RAM | Savings | tok&#x2F;s | |-------|------|---------|---------|-------| | gpt-oss:120b | 58.5GB | 19.1GB | -67% | 3.84 | | qwen3:30b-a3b | 18.0GB | 10.4GB | -42% | 19.7 | | gemma4:26b-a4b | 16.2GB | 10.6GB | -35% | 11.5 | | GLM-4.7-Flash | 19.0GB | 13.3GB | -30% | 12.4 |<p>Output is SHA-256 bit-identical to full-RAM runs. Verified.<p>How it works (~50 lines of C++ in llama.cpp): 1. Profile MoE expert routing frequencies. 2. mmap the entire model from disk. 3. mlock only the hottest experts into physical RAM. 4. posix_fadvise &#x2F; prefetch cold experts from NVMe before they&#x27;re needed.<p>Boundary condition: if Disk &gt; RAM, pinning gives +45% speedup (gpt-oss: 2.64 → 3.84 tok&#x2F;s). If model fits in RAM, pinning adds zero overhead.<p>Tested on: Linux native, WSL, Windows native (VirtualLock).<p>Repo: https:&#x2F;&#x2F;github.com&#x2F;LozzKappa&#x2F;hotpin-llm Paper (PDF + LaTeX) in the repo – arXiv submission pending endorsement in cs.LG.<p>I&#x27;m looking for: 1. An arXiv endorser (cs.LG) – if you&#x27;re a researcher and this work interests you, please reach out. 2. Feedback from anyone who tests it on their hardware.<p>The technique is simple, lossless, and works today. Test it and tell me your benchmarks.<p>Thanks for reading.