展示 HN:将 Web 应用逆向工程为代理工具

12 分•作者: pancomplex•3 个月前
各位 HN 的朋友们!我们构建了一个浏览器内嵌代理,它运行在一个已认证的 Web 应用中,能够观察该应用如何调用自身的 API,并自动将这些调用转化为代理工具。你可以将其理解为一个自动生成的 MCP 服务器,并且会随着宿主应用的变更而自我更新。 最终,这将是一个技能娴熟的 AI 助手,能够以最小的努力深度集成到任何产品中(而不仅仅是聊天和 RAG)。 请查看下面的简短演示,它们展示了该代理在您可能熟悉的一些软件中的应用: * Jira: https://demo.frigade.com/hn?skill=jira * Spotify: https://demo.frigade.com/hn?skill=spotify * Hacker News (哈哈): https://demo.frigade.com/hn?skill=hackernews * 完整演示: https://demo.frigade.com/hn?skill=full-demo 正如您在示例中看到的,您可以比通常通过点击操作做得更多(也更快)。而且我们甚至没有接触过这些产品的源代码! 为什么要做这个? 理想情况下,每个应用程序都应该有一个 MCP 服务器或一个易于理解的 API,供 AI 代理从中获取信息。实际上,我们发现即使是非常现代化的软件,也往往拥有一个由混乱的 API 和服务组成的“蜘蛛网”,AI 代理根本无法开箱即用。安全性也成为了一个巨大的问题,因为应用程序在如何保护端点方面有不同的(通常是自制的)标准(JWT/Cookie/两者的混合)。最后,让一个实际的浏览器代理代表用户进入并使用应用程序(即计算机使用)是极其脆弱、缓慢且会消耗大量 token 的。 我们利用了我们现有的浏览器代理,它已经能够使用和学习已认证的应用程序,并增加了一个额外的步骤,可以自动将应用程序的已认证 API 转化为“食谱”。一个食谱是以下内容的组合: * API 端点 + 方法 * 认证方法(以及如何检索刷新认证 token/Cookie) * 响应模式 * 输入模式(用于 POST/PUT) * 对工具功能的易于理解的描述 将所有这些整合在一起,它们就成为了 LLM 的可重用工具,而且无需编写或维护任何代码。即使 API 发生变化,我们的代理也能发现并用更新后的版本替换工具的食谱。 通过这种方式,为 AI 代理添加工具变得超级简单: * 我们的代理学习应用程序并构建食谱。 * 应用程序所有者从我们的仪表板中启用已发现的工具。 * 代理现在可以直接在应用程序内代表用户执行操作。例如,说“邀请我的队友加入我的工作区”将安全地调用现有的用户邀请 API 端点,而无需通过第三方进行代理或中继。 当然,在尝试这样做时会遇到大量的边缘情况——尽管存在许多“标准”,但每个应用程序本质上都是不同的。有趣的事实是:在标准化食谱方面,GraphQL 是迄今为止最难处理的 API。 期待您的反馈/评论!
查看原文
Hey HN! We built a browser-based agent that runs inside an authenticated web app, watches how the app calls its own APIs, and automatically turns those into agent tools. You can think of it as an auto-generated MCP server that self-updates as the host app changes.<p>The result is a skilled AI assistant that actually integrates deeply with any product (not just chat and RAG) with minimal effort.<p>Check out these short demos below that show the agent in software you&#x27;re probably familiar with:<p>- Jira: <a href="https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=jira">https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=jira</a><p>- Spotify: <a href="https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=spotify">https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=spotify</a><p>- Hacker News (lol): <a href="https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=hackernews">https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=hackernews</a><p>- Full Demo: <a href="https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=full-demo">https:&#x2F;&#x2F;demo.frigade.com&#x2F;hn?skill=full-demo</a><p>As you can see in the examples, you can do way more (and faster) than what you normally would be able to via point and click. And we never even touched the source code of these products!<p>Why do this?<p>In an ideal world, every application has an MCP server or an easily-digestible API available for AI agents to feed from. In practice, we found that even very modern software tends to have a spider web of confusing APIs and services that AI agents simply cannot use out of the box. Security also becomes a huge issue as applications have different (often homebrewed) standards for how endpoints are secured (JWTs&#x2F;cookies&#x2F;mix of both). Finally, having an actual browser agent go in and use the application on behalf of the user (i.e. computer-use), is simply too brittle, slow, and burns a lot of tokens.<p>We took our existing browser agent that’s already trained to use and learn authenticated applications, and added an extra step that automatically turns the app’s authenticated APIs into &quot;recipes&quot;. A recipe is a mix of the following:<p>- API endpoint + method<p>- Authentication method (and how to retrieve refresh auth tokens&#x2F;cookies)<p>- Response schema<p>- Input schema (for POST&#x2F;PUT)<p>- Human readable description of what the tool does<p>Putting it all together, these become reusable tools for LLMs, all without writing or maintaining any code. Even if the APIs change our agent figures this out and replaces the recipe for the tool with the updated version.<p>Adding tools to an AI agent becomes super simple this way:<p>- Our agent trains on the app and builds the recipes<p>- The app owner enables discovered tools from our dashboard<p>- The agent can now take actions on the user’s behalf directly inside the application. For instance, saying something like &quot;invite my teammate to my workspace&quot; would securely call the existing API endpoint for inviting users without proxying or relaying through a third party.<p>Of course, there&#x27;s a ton of edge cases you run into when you try to do this - every application is intrinsically different despite how many &quot;standards&quot; exist. Fun fact: graphql was by far the worst API to work with in standardizing the recipes.<p>Looking forward to your feedback&#x2F;comments!