Show HN:BlazeRules – 用于流式数据的 YAML 规则引擎,每秒可处理高达 500 万条记录

1 分•作者: jspuri•大约 2 个月前
<a href="https:&#x2F;&#x2F;github.com&#x2F;purijs&#x2F;blazerules" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;purijs&#x2F;blazerules</a><p>我最初想用 C++ 编写一个亚毫秒级的日志解析器,但后来演变成了一个可嵌入的决策引擎,可以对传入数据运行 YAML 定义的规则。 这些规则通过先重新投影为列式格式(如果尚未是列式格式)来以向量化格式在传入数据上执行。根据载荷大小和规则的复杂性,性能从每秒 200,000 条记录到每秒超过一百万条记录不等,吞吐量平均约为每秒 200 MiB 到每秒 3 GiB。</p><p>规则也可以是 SQL 表达式,或者 ONNX 模型(数值型)、窗口操作以及支持的许多其他操作。</p><p>它与 DuckDB 类似,但适用于流式数据和即时决策。</p>
查看原文
<a href="https:&#x2F;&#x2F;github.com&#x2F;purijs&#x2F;blazerules" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;purijs&#x2F;blazerules</a><p>I initially wanted to make a sub-millisecond log parser in C++ but that blew into a embeddable decision engine, that can run YAML defined rules on incoming data. The rules are executed in a vectorized format on incoming data by reprojecting into a columnar format first, if it&#x27;s not already. Depending on the payload size and rules complexity, the performance goes from 200K records&#x2F;s to more than million records&#x2F;sec, in terms of througput this would be around 200 MiB&#x2F;s to 3 GiB&#x2F;s on average.<p>Rules can be sql expressions too, or onnx models (numeric), window ops and quite a few more operations are supported.<p>It&#x27;s comparable to DuckDB but for streaming data and on the fly decisions.