Launch HN:Parsewise (YC P25) – 通过 API 实现跨文档推理
5 分•作者: gergelycsegzi•3 个月前
大家好,我是 Parsewise 的创始人 Greg 和 Max。
Parsewise 能将一堆非结构化数据转化为符合模式(schema)要求的数据,并保留跨文档解析值的溯源信息。
想象一下,你给 Claude 一堆文件,要求它输出 CSV 或 JSON 格式。如果你尝试过,就会知道这其中存在系统限制(文件数量、输入类型、成本、延迟),以及一个面向用户的挑战:没有快速验证结果的方法。我们解决了这两者。
我们帮助技术团队简化非结构化数据 ETL,并让业务专家参与定义和即时验证。
以下是一段展示一些用例的视频:[https://www.youtube.com/watch?v=dbRllnnh47w](https://www.youtube.com/watch?v=dbRllnnh47w)
用一位早期用户的话来说 Parsewise:
“我需要从保险单 PDF、转录的电话录音、电子邮件等中提取信息。我寻找的不是那种只能逐个数据点、逐页提取到结构良好、定义明确的模式的东西,而是更具智能代理性质的工具,能够理解信息可能分散在不同文档中,并能基于此进行推理以提取所需内容。”
我们创立这家公司是基于十年来在复杂数据转换和数据分析/综合方面的经验(以及痛点)。Greg 曾在 Palantir 构建过传统的 ETL 和实施过 AI 工作流。Max 在 Bain 曾对金融行业的复杂数据进行分析,这与我们许多客户的情况类似。
Parsewise 的工作方式是接收一堆数据(例如数百或数千个 PDF、Excel 文件等),然后输出符合模式要求的数据,其中每一个值都可以追溯到原始文档中词级别的引用。我们为 API 客户提供在其自有应用程序中展示溯源信息的方式,或者他们可以使用我们的平台进行内部操作。
在数据处理的核心,我们拥有自我改进的代理(agent)定义。它们定义了可接受的数据源、解析或合并值的逻辑,以及向最终用户突出显示不确定性的规则。
底层技术支持模型和云无关部署,并可在私有网络中部署。我们在视觉推理方面使用 Gemini 模型取得了最佳效果,在最强大的基于事实推理的基准测试(Databricks OfficeQA)上达到了最先进水平(优于 Claude Fable)[将附上我们博客文章的链接]。
值得注意的是,我们更侧重于“人工验证”(human harness),而不是模型本身(model harness),因为我们发现实际应用中的主要障碍是可验证性。这意味着我们要优化信任结果所需的时间和点击次数。
我们使用 vLLMs 进行解析,然后使用小型模型进行高效的大规模穷尽搜索。与 RAG 不同,我们不进行采样;相反,我们穷尽地查找给定查询的所有相关值。我们使用大型模型来处理解析结果的决策以及向用户标记不一致之处。
这种穷尽性和明确的值来源是我们的平台独有的,它超越了许多现有提供商仅覆盖数据解析第一步的功能。
我们非常欢迎开发者和技术爱好者尝试 Parsewise 来解决你们复杂的文档挑战。我们有很多关于如何扩展产品和使其变得更好的想法,但也非常期待社区的反馈和建议!
查看原文
Hi all, it’s Greg and Max, founders of Parsewise here<p>Parsewise transforms a bucket of unstructured data into schema compliant data retaining lineage for values resolved across documents.
Imagine giving Claude a bunch of files and asking for a CSV or JSON output. If you have tried this, you know both the system limitations (number of files, type of inputs, cost, latency) but also the human-facing challenge of having no way to validate the results quickly. We solve both.
We help tech teams simplify their unstructured data ETL, and loop in business experts for the definitions and for instant validation.<p>Here is a video with a few use cases: <a href="https://www.youtube.com/watch?v=dbRllnnh47w" rel="nofollow">https://www.youtube.com/watch?v=dbRllnnh47w</a><p>Parsewise in the words of someone coming to us:
”I need to extract information from insurance policy PDFs, phone calls that have been transcribed, emails, etc. I am NOT looking for something that would just extract data point by data point, page by page into a structured well-defined schema but more something more agentic that can understand that information might be across documents and that it should reason over what to extract.”<p>We started the company based on a decade of experience (and pain) in complex data transformation and data analysis / synthesis. Greg was building both classical ETL and implemented AI workflows at Palantir. At Bain, Max did highly complex data analysis in the financial sector, similar to many of our customers.<p>Parsewise works by taking in a bucket of data (think hundreds or thousands of pdfs, excels etc.), and outputting schema compliant data where every single value is traceable down to word level citations across multiple documents in the bucket. We provide API customers with ways to show the lineage in their own applications, or they can use our platform for internal operations.
At the core of the data processing we have self-improving agent definitions. They define the acceptable sources, the logic for resolving or combining values, and the rule for highlighting uncertainty to the end user.<p>The underlying tech is model and cloud agnostic and can be deployed in private networks. We have seen the best results with Gemini models for visual reasoning, achieving SOTA (beating Claude Fable) on the strongest grounded reasoning benchmark we have found (Databricks OfficeQA) [will include link to our blog post].
Notably, we focused more on the “human harness” rather than the model harness, leaning into the actual friction we saw in uptake, which is around verifiability. That means optimizing the time and clicks required to trust the outcomes.
We use vLLMs for parsing, and then we use small models for efficient large scale exhaustive search. Unlike RAG, we do not sample; instead, we exhaustively find all relevant values for a given query. We use larger models for decision making around resolutions and flagging inconsistencies to users.<p>This exhaustiveness and explicit value sourcing is unique to our platform, and it goes beyond the first step of data parsing that many existing providers cover.<p>We would love to welcome builders and tinkerers to try Parsewise on your complex document challenges. We have a ton of ideas on how we can expand the product and make it better, but would appreciate feedback and ideas from the community!