Show HN:Morph Reflexes – 用于智能体轨迹的多头分类器

1 分•作者: bhaktatejas922•3 个月前
生产代理最常见的故障是行为方面的:循环、推理泄露、用户沮丧等。使用 GPT 或 Sonnet 等前沿模型来判断每一个回合的成本太高且速度太慢,无法大规模运行。 工作原理: 我们使用具有混合注意力的现代 LLM,并移除解码步骤。我们构建了一个推理引擎,允许预填充计算从一个反射到另一个反射重复使用 99%,这在精神上类似于 2019 年的 BERT/HYDRA + 旧的多头技术。 我们采用了相同的宏观理念,并付出了艰苦的努力,使其能够与现代架构和注意力机制协同工作。在此基础上,我们可以在 30 毫秒内完成推理,并在 90 毫秒内处理完整个请求。无论您运行 4 个反射还是 100 个反射,额外的开销都不到 2 毫秒。 为什么优化这一点很重要? 即使您是一个中等规模的初创公司,您也会处理数万次代理运行和数百万个回合。如果您想跟踪用户沮丧率随时间的变化,前沿 LLM 作为裁判的方案是无法扩展的。 我在特斯拉构建了一个类似的堆栈。当 ML 工程师需要跨 PB 级数据采样诸如 `is_camera_obfuscated=true` 等信号以及其他 200 种信号时,您需要 1) 快速启动它们 2) 高效地大规模运行。 它不是什么: 仪表板。根据我的经验,99% 的仪表板都不会被使用。这个工具纯粹是基于 API 的,专为希望自行跟踪代理行为、触发自己的警报并在此基础上进行构建的开发者而设计。 您可以在我们的仪表板中训练自定义反射,然后让它在生产环境中自我改进:https://www.morphllm.com/dashboard/reflex 文档:https://docs.morphllm.com/sdk/components/reflexes/index 我很想听听那些在生产环境中运行代理的人的反馈:您希望能够跟踪哪些类型的信号,覆盖 100% 的回合? 简而言之:来自代理跟踪的语义信号,速度极快,通过 API 实现成本低廉。
查看原文
The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale.<p>How it works:<p>We use a modern LLM with hybrid attention and remove the decode step. We built an inference engine that lets prefill compute be 99% reused from reflex to reflex, similar in spirit to older 2019-era BERT&#x2F;HYDRA + older multiple-head techniques.<p>We took the same high-level idea and did the hard work to make it work with a modern architecture and attention. On it, we can run inference in under 30ms and serve the full request in under 90ms. If you run 4 reflexes or 100, the extra overhead is less than 2ms.<p>Why does optimizing this matter?<p>If you’re even a medium-sized startup, you’re dealing with tens of thousands of agent runs and millions of turns. If you want to track things like user frustration rates over time, frontier LLM-as-judge does not scale.<p>I built a similar stack at Tesla. When ML engineers needed to sample data across petabytes for signals like `is_camera_obfuscated=true`, along with 200 other things, you need to 1) spin them up quickly 2) run at scale efficiently<p>What it is not:<p>A dashboard. In my experience, 99% of dashboards go unused. This is purely API-based and made for devs who want to track agent behavior themselves and trigger their own alerts and build on it.<p>You can vibetrain a custom reflex in our dashboard, and then let it self improve in production: <a href="https:&#x2F;&#x2F;www.morphllm.com&#x2F;dashboard&#x2F;reflex">https:&#x2F;&#x2F;www.morphllm.com&#x2F;dashboard&#x2F;reflex</a><p>Docs: <a href="https:&#x2F;&#x2F;docs.morphllm.com&#x2F;sdk&#x2F;components&#x2F;reflexes&#x2F;index">https:&#x2F;&#x2F;docs.morphllm.com&#x2F;sdk&#x2F;components&#x2F;reflexes&#x2F;index</a><p>I’d love feedback from people running agents in prod: what sorts of things do you wish you could track over time across 100% of turns?<p>TLDR: semantic signals from agent traces, super fast, cheap via API