Show HN:Cactus v2 – 设备端 AI,支持云端回退

1 分•作者: rshemet•3 个月前
大家好,我是 Cactus 的 Roman 和 Henry(<a href="https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus</a>)。 我们刚刚发布了我们设备端推理平台的重大升级: * 内置基于模型置信度的路由,可将推理任务转交给云端。 * 支持任何 PyTorch 模型的转换器。 * 无损 4 位量化(评估结果请参见我们的 GitHub README)。 * 支持兼容设备的 GPU 加速(首批支持 Apple Metal)。 * 极低的 RAM 占用。 * 可在任何 Arm 设备上运行:iOS、Android、Mac、DGX Spark、Raspberry Pi 等。 总而言之,一个 Gemma 4 E2B 级别的模型在 M5 Max 上可达到 169 token/秒 的速度,占用 2.7GB 磁盘空间且与 FP16 相比无精度损失,使用 1.3GB RAM,并在需要时向云端模型寻求帮助。 我们十八个月前开始解决的问题是:推理引擎大多为数据中心设计,但消费级硬件的物理特性不同:RAM 需要与操作系统共享,会受到热节流限制,并且同一模型在不同硬件上的表现也不同。 因此,我们为资源受限的设备从头开始编写了一个运行时。自那时以来,Cactus 已发展到每周处理数百万次推理,并拥有数万名月活跃开发者。 我们在生产应用中部署 Cactus 的最大体会是,虽然本地模型可以处理 90% 的工作负载,但那 10% 的差距意味着它们仍然不够生产就绪。我们的用户通常通过构建自定义的云端回退逻辑来解决这个问题。 Cactus v2 解决了这个问题: 我们实现云端回退的方法是,在模型的权重中植入一个探针,该探针读取模型的内部激活并发出置信度信号。这样,路由就发生在模型内部,而不是由一个位于其前方的提示分类器来完成。我们认为这对于多轮智能代理工作至关重要,模型应该知道哪些轮次足够简单可以在本地处理,哪些轮次比较困难——然后将其转交给云端。本次发布支持 Gemma-4 E2B 的单轮路由,可配置为向任意端点(Gemini、Claude、OpenAI 兼容端点或您自己的端点)进行升级。 我们的下一个目标是为多轮智能代理工作开发混合原生模型。这是真正尚未解决的问题。我们目前的权重内探针方法已显示出有希望的结果,并计划很快发布首批模型变体。 混合变体是 Gemma 的衍生品,根据 HF 卡上注明的 Gemma 条款发布。 除了混合路由,该运行时还具备最先进的量化技术,能在 4 位精度下实现无损量化,通过内存映射权重来减少 RAM 占用,并支持跨平台运行,提供 Python、Rust、React Native、Swift 和 Kotlin 绑定。 免责声明:Cactus 以源代码可用形式分发——个人使用和小型公司免费;超出此范围的商业使用需付费(类似 Docker 的许可模式)。 您可以通过我们的 GitHub (<a href="https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus</a>) 或通过 `brew install cactus-compute&#x2F;cactus&#x2F;cactus` 开始使用。
查看原文
Hi HN, Roman and Henry here from Cactus (<a href="https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus</a>).<p>We just shipped the biggest upgrade to our on-device inference platform:<p>- Built-in model confidence-based routing to hand off inference runs to the cloud - Converter for any PyTorch model - Lossless 4-bit quantization (evals on our GitHub README) - GPU acceleration on compatible devices (starting with Apple Metal) - Minimal RAM footprint - Runs on any Arm device: iOS, Android, Mac, DGX Spark, Raspberry Pi, and more<p>All in, a Gemma 4 E2B class model runs at 169 tok&#x2F;sec on M5 Max, takes 2.7GB disk space with no accuracy degradation from FP16, uses 1.3GB of RAM, and requests help from cloud models when needed.<p>The problem we started with eighteen months ago: inference engines are built for datacenters, but consumer hardware has different physics: you share RAM with the OS, you get thermally throttled, and the same model behaves differently on different hardware.<p>So we wrote a runtime from scratch for resource-constrained devices. Since then, Cactus has grown to process millions of weekly inference runs, and tens of thousands of monthly active developers.<p>Our biggest learning from deploying Cactus in production apps is that while local models can handle 90% of workloads, that 10% gap means they&#x27;re still not production-ready. Our users&#x27; fix was to build custom cloud fallback logic.<p>Cactus v2 fixes that:<p>Our approach to cloud fallback is to post-train a probe into the model&#x27;s weights that reads its internal activations and emits a confidence signal. This way, the routing happens inside the model rather than in a prompt classifier sitting in front of it. We believe this is critical for multi-turn agentic work, where the model should know which turns are easy enough to be handled locally, and which are hard - and get handed off to the cloud. What ships today is single-turn routing for Gemma-4 E2B against a configurable escalation endpoint (Gemini, Claude, OpenAI-compatible, or your own endpoint).<p>Our next target is hybrid-native models for multi-turn agentic work. This is the genuinely unsolved problem. Our current probe-in-the-weights is showing promising results and we look to release the first model variants soon.<p>The hybrid variants are Gemma derivatives, released under the Gemma terms noted on the HF cards.<p>In addition to the hybrid routing, the runtime has SOTA quantization, which is lossless at 4bit, memory maps weights to decrease RAM footprint and runs cross-platform, with Python, Rust, React Native, Swift, and Kotlin bindings.<p>Disclaimer: Cactus is distributed as source-available - free for personal use and small companies; commercial license above that (Docker-style license).<p>You can get started on our GitHub: <a href="https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;cactus-compute&#x2F;cactus</a> or by `brew install cactus-compute&#x2F;cactus&#x2F;cactus`.