展示 HN:Hands-Rust MCP/CLI,可查看 Windows 桌面并点击真实的 Chrome
1 分•作者: ryan-b•大约 1 个月前
我构建 Hands 的初衷是希望有一个编码代理能够像我一样使用这台 Windows PC 和真实的 Chrome 配置文件:查看屏幕、移动真实鼠标、输入文字、点击,而无需将 Chrome 转变为自动化浏览器。
它是一个 Rust MCP/CLI。一个“线束”(如 Grok、Codex、Claude Code、OpenCode 等)会调用 observe、click、type、scroll 等工具。Observe 是一个屏幕截图路径加上一个小的元素列表(UIA + 可选的 Chrome DOM ID)。Click 是通过 Bézier 曲线在操作系统上使用 SendInput 实现的,而不是通过 Chrome DevTools 的点击。
没有 Playwright,没有 Puppeteer,也没有远程调试端口。日常使用的 Chrome 会正常启动,不会添加额外的标志,或者如果它已经打开,则会附加到现有会话。大多数依赖 CDP/自动化标志的网站无法检测到这一点。它们仍然可以看到注入的输入(LLMHF_INJECTED)。
一个微小的未打包的 Chrome 扩展程序可以融合页面结构(chr: ID、列表卡片),这样模型就不会仅仅依赖像素进行猜测。侧载是手动进行的。如果服务工作线程不活跃,融合就会失效;需要重新加载卡片。
它的优点:
* 适用于您自己的桌面上的个人研究。“在 cars.com 上查找 Camry”、“阅读页面”、“输入 ZIP 码”、“关闭 cookie 横幅”。
它的局限性:
* 不是沙箱。它可以点击屏幕上的任何内容,包括结账和“一键申请”。
* “确认后付款”是二元分类中的尽力而为,并非保证。来自屏幕截图/DOM 的提示注入是真实存在的;二进制文件将该文本视为不受信任的,而模型可能不会。
* 在日常 Chrome 中不是 CAPTCHA 求解器。可见的尝试两次后,它会放弃并等待谜题消失。
* 仅限 Windows。
* 安装步骤:构建 exe 文件,注册一个原生消息传递主机,侧载扩展程序,将 MCP 客户端指向 Hands MCP。README 文件是操作指南。缺少 API 密钥不会导致构建失败;do_task 是可选的。
* 日志位于 %LOCALAPPDATA%\hands\logs\ 下。扩展程序请求 `<all_urls>` 权限,以便能够映射您正在查看的标签页。
代码库:<a href="https://github.com/Ryan-AI-Studios/hands" rel="nofollow">https://github.com/Ryan-AI-Studios/hands</a> (MIT)
我很乐意解答关于 observe/fusion/the fence 工作原理的问题。如果您尝试使用它,Pause/Break 键是终止开关。
查看原文
I built Hands because I wanted a coding agent to use this Windows PC and a real Chrome profile the way I do: look at the screen, move the real mouse, type, click , without turning Chrome into an automation browser.<p>It is a Rust MCP/CLI. A harness (Grok, Codex, Claude Code, OpenCode, etc.) calls tools like observe, click, type, scroll. Observe is a screenshot path plus a small element list (UIA + optional Chrome DOM ids). Click is OS SendInput on a Bézier path, not a Chrome DevTools click.<p>There is no Playwright, no Puppeteer, no remote debugging port. Daily Chrome is launched with no extra flags, or attached if it’s already open. Sites that key on CDP/automation flags mostly don’t see that. They can still see injected input (LLMHF_INJECTED).<p>A tiny unpacked Chrome extension can fuse page structure (chr: ids, listing cards) so the model isn’t guessing from pixels. Sideload is manual. Fusion dies if the service worker goes inactive; reload the card.<p>What it is good for: personal research on your own desk. “Find a Camry on cars.com,” read a page, fill a ZIP, dismiss a cookie banner.<p>What it is not:
• Not a sandbox. It can click whatever is on screen, including checkout and Easy Apply.
• Confirm-before-money is best-effort classification in the binary, not a guarantee. Prompt injection from the screenshot/DOM is real; the binary treats that text as untrusted, the model might not.
• Not a CAPTCHA solver on daily Chrome. Two visible tries, then it yields and waits for the puzzle to go away.
• Windows only.
• Install is: build the exe, register a native-messaging host, sideload the extension, point an MCP client at hands mcp. README is the runbook. Missing an API key does not fail the build; do_task is optional.
• Logs live under %LOCALAPPDATA%\hands\logs\. The extension asks for <all_urls> so it can map the tab you’re looking at.<p>Repo: <a href="https://github.com/Ryan-AI-Studios/hands" rel="nofollow">https://github.com/Ryan-AI-Studios/hands</a> (MIT)<p>Happy to answer how observe/fusion/the fence work. If you try it, Pause/Break is the kill switch.